LLM System Design Interview, issue 66, Sep 10, 2026

The Optimizer Scaling Trap

Staff Research Engineer interview at Anthropic, and the interviewer asks:

Your team wants to swap AdamW for Muon on the next frontier run because it wins on small-scale benchmarks. What two axes must their ablation cover before you sign off?

Don’t say: Run it at 3 model sizes and check the loss curve holds.

Why a 2x small-scale speedup quietly evaporates at frontier compute, and the hidden token-to-parameter trap that ruins multi-million-dollar training runs.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.