LLM System Design Interview, issue 66, Sep 10, 2026
The Optimizer Scaling Trap
Staff Research Engineer interview at Anthropic, and the interviewer asks:
“Your team wants to swap AdamW for Muon on the next frontier run because it wins on small-scale benchmarks. What two axes must their ablation cover before you sign off?”
Don’t say: “Run it at 3 model sizes and check the loss curve holds.”
Why a 2x small-scale speedup quietly evaporates at frontier compute, and the hidden token-to-parameter trap that ruins multi-million-dollar training runs.
The full answer, with the mechanism and the arithmetic, is free on Substack.