Advanced NLP Interview Questions, issue 2, Dec 9, 2025

The Gradient Shockwave Trap

Senior ML Engineer interview at NVIDIA, and the interviewer asks:

You attach a new, random linear head to a pre-trained Transformer. Do you unfreeze all layers and start backprop immediately?

How a single random layer can silently erase millions of dollars of pretraining, and how to stop it.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.