LLM Inference Interview Questions, issue 17, Aug 17, 2026
The Reasoning Budget Trap
Staff AI Engineer interview at Anthropic, and the interviewer asks:
“Your agent scores 12 points higher with reasoning effort set to high, and 40% of users abandon the session before it finishes. How do you decide where to spend that reasoning budget?”
Don’t say: “I’d use a smaller model for easy tasks and route hard ones to the big model.”
How maximum thinking time kills user retention by minute four, and the "escalate-on-failure" trick that buys 58% success rates without the 11-minute latency tax.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.