LLM System Design Interview, issue 63, Sep 7, 2026
The Intermediate Loss Trap
Senior ML Engineer interview at Anthropic, and the interviewer asks:
“You switched your scaling law sweeps from cosine to WSD (warmup-stable-decay). Two weeks in, the intermediate loss curves look strictly worse than the cosine baselines and leadership wants to roll back. What do you tell them?”
Don’t say: “WSD is a bit worse than cosine, we should revert.”
The hidden reason your model looks strictly worse for 80% of training, and how holding your nerve turns quadratic sweep costs into near-free scaling data points.
The full answer, with the mechanism and the arithmetic, is free on Substack.