Advanced NLP Interview Questions, issue 1, Dec 9, 2025

The Learning Rate Warm-Up Trap

Senior ML Engineer interview at Google DeepMind, and the interviewer asks:

โ€œWe are training a ๐˜›๐˜ณ๐˜ข๐˜ฏ๐˜ด๐˜ง๐˜ฐ๐˜ณ๐˜ฎ๐˜ฆ๐˜ณ from scratch using ๐˜ˆ๐˜ฅ๐˜ข๐˜ฎ. We set a constant Learning Rate of 1e-3. Predict the first 1000 steps.โ€

Donโ€™t say: โ€œIt converges. ๐˜ˆ๐˜ฅ๐˜ข๐˜ฎ is an adaptive optimizer, it adjusts per-parameter learning rates automatically. 1e-3 is a standard default. It might be noisy at first, but the loss will go down.โ€

The hidden physics of why Transformers can't handle a 1e-3 LR at initialization.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.