Advanced NLP Interview Questions, issue 7, Dec 13, 2025

The Exploding Gradient Trap

Senior ML Engineer interview at Meta, and the interviewer asks:

You’re training a 7B parameter Llama-style model. In the first 1000 steps, your gradients start oscillating wildly and the loss spikes. How do you fix it?

Why touching the learning rate feels right… and silently destroys large-model training.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.