LLM System Design Interview, issue 62, Sep 6, 2026
The muP Transfer Trap
Senior AI Research Engineer interview at Anthropic, and the interviewer asks:
“Your team burned three weeks sweeping learning rates at every rung of the scaling ladder before the 7B run. A colleague says muP would have let you tune once at 100M and transfer for free. Were they right?”
Don’t say: “Yes, muP makes the optimal learning rate width-invariant, so you tune small and transfer up.”
Why relying on "free" learning rate transfer can silently degrade your 7B run, and the hidden architectural invariants you must audit before betting the cluster.
The full answer, with the mechanism and the arithmetic, is free on Substack.