LLM System Design Interview, issue 57, Sep 1, 2026

The Expressiveness Trap

Staff ML Engineer interview at Anthropic, and the interviewer asks:

Your model is 80 layers deep and your infra team is furious. Why does almost every production LLM land near a 100:1 width-to-depth ratio?

Don’t say: Deeper models are more expressive, so it’s a capacity tradeoff.

Why fighting for a 0.5% loss improvement via deeper models is a massive systems mistake and why the 100:1 width ratio is actually an MFU survival tactic.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.