Advanced NLP Interview Questions, issue 2, Dec 9, 2025
The Gradient Shockwave Trap
Senior ML Engineer interview at NVIDIA, and the interviewer asks:
“You attach a new, random linear head to a pre-trained Transformer. Do you unfreeze all layers and start backprop immediately?”
How a single random layer can silently erase millions of dollars of pretraining, and how to stop it.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.