Advanced Deep Learning Interview Questions, issue 9, Mar 30, 2026

The Local Minimum Trap

Senior ML Engineer interview at Google DeepMind, and the interviewer asks:

You’re training a massive 50-billion parameter MLP on a cluster of H100 GPUs. Your monitoring tool shows the gradient norm has hit absolute zero, but your loss is still unacceptably high. What just happened?

A flat gradient isn’t success, it’s often a saddle where curvature, not slope, determines whether your model is actually stuck.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.