Advanced NLP Interview Questions, issue 12, Dec 18, 2025
The Optimizer State Memory Trap
Senior Machine Learning Engineer interview at Google DeepMind, and the interviewer asks:
โYou just switched a 7B parameter training run from SGD to Adam to speed up convergence. The model size is identical, but the cluster immediately crashes with a ๐๐๐๐ ๐๐ถ๐ต-๐๐ง-๐๐ฆ๐ฎ๐ฐ๐ณ๐บ (๐๐๐) error. Why?โ
Donโt say: โAdam is computationally more expensive so it uses more memory,โ
Why switching from SGD to Adam can instantly triple your VRAM usage and crash a 7B training run.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.