Advanced NLP Interview Questions, issue 12, Dec 18, 2025

The Optimizer State Memory Trap

Senior Machine Learning Engineer interview at Google DeepMind, and the interviewer asks:

โ€œYou just switched a 7B parameter training run from SGD to Adam to speed up convergence. The model size is identical, but the cluster immediately crashes with a ๐˜Š๐˜œ๐˜‹๐˜ˆ ๐˜–๐˜ถ๐˜ต-๐˜–๐˜ง-๐˜”๐˜ฆ๐˜ฎ๐˜ฐ๐˜ณ๐˜บ (๐˜–๐˜–๐˜”) error. Why?โ€

Donโ€™t say: โ€œAdam is computationally more expensive so it uses more memory,โ€

Why switching from SGD to Adam can instantly triple your VRAM usage and crash a 7B training run.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.