Advanced Deep Learning Interview Questions, issue 9, Mar 30, 2026
The Local Minimum Trap
Senior ML Engineer interview at Google DeepMind, and the interviewer asks:
“You’re training a massive 50-billion parameter MLP on a cluster of H100 GPUs. Your monitoring tool shows the gradient norm has hit absolute zero, but your loss is still unacceptably high. What just happened?”
A flat gradient isn’t success, it’s often a saddle where curvature, not slope, determines whether your model is actually stuck.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.