LLM System Design Interview, issue 63, Sep 7, 2026

The Intermediate Loss Trap

Senior ML Engineer interview at Anthropic, and the interviewer asks:

You switched your scaling law sweeps from cosine to WSD (warmup-stable-decay). Two weeks in, the intermediate loss curves look strictly worse than the cosine baselines and leadership wants to roll back. What do you tell them?

Don’t say: WSD is a bit worse than cosine, we should revert.

The hidden reason your model looks strictly worse for 80% of training, and how holding your nerve turns quadratic sweep costs into near-free scaling data points.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.