LLM System Design Interview, issue 62, Sep 6, 2026

The muP Transfer Trap

Senior AI Research Engineer interview at Anthropic, and the interviewer asks:

Your team burned three weeks sweeping learning rates at every rung of the scaling ladder before the 7B run. A colleague says muP would have let you tune once at 100M and transfer for free. Were they right?

Don’t say: Yes, muP makes the optimal learning rate width-invariant, so you tune small and transfer up.

Why relying on "free" learning rate transfer can silently degrade your 7B run, and the hidden architectural invariants you must audit before betting the cluster.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.