Advanced NLP Interview Questions, issue 1, Dec 9, 2025
The Learning Rate Warm-Up Trap
Senior ML Engineer interview at Google DeepMind, and the interviewer asks:
โWe are training a ๐๐ณ๐ข๐ฏ๐ด๐ง๐ฐ๐ณ๐ฎ๐ฆ๐ณ from scratch using ๐๐ฅ๐ข๐ฎ. We set a constant Learning Rate of 1e-3. Predict the first 1000 steps.โ
Donโt say: โIt converges. ๐๐ฅ๐ข๐ฎ is an adaptive optimizer, it adjusts per-parameter learning rates automatically. 1e-3 is a standard default. It might be noisy at first, but the loss will go down.โ
The hidden physics of why Transformers can't handle a 1e-3 LR at initialization.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.