Advanced NLP Interview Questions, issue 7, Dec 13, 2025
The Exploding Gradient Trap
Senior ML Engineer interview at Meta, and the interviewer asks:
“You’re training a 7B parameter Llama-style model. In the first 1000 steps, your gradients start oscillating wildly and the loss spikes. How do you fix it?”
Why touching the learning rate feels right… and silently destroys large-model training.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.