LLM System Design Interview, issue 57, Sep 1, 2026
The Expressiveness Trap
Staff ML Engineer interview at Anthropic, and the interviewer asks:
“Your model is 80 layers deep and your infra team is furious. Why does almost every production LLM land near a 100:1 width-to-depth ratio?”
Don’t say: “Deeper models are more expressive, so it’s a capacity tradeoff.”
Why fighting for a 0.5% loss improvement via deeper models is a massive systems mistake and why the 100:1 width ratio is actually an MFU survival tactic.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.