LLM System Design Interview, issue 75, Sep 19, 2026
The Thinking Fusion Trap
Staff AI Engineer interview at Anthropic, and the interviewer asks:
“Product wants one model that flips between instant answers and long reasoning with a prompt tag, so we only run one deployment. Reasoning benchmarks dropped 2 points. Do you ship it?”
Don’t say: “Two points is within noise, ship it, the infra savings are worth it.”
The hidden gradient conflict when forcing one set of weights to both terminate fast and think deep, and why "saving serving costs" quietly destroys your highest-value reasoning.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.