LLM Inference Interview Questions, issue 19, Aug 19, 2026
The 10x Sampling Trap
Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
“You 10x’d your sampling budget on a hard reasoning task and the solve rate barely moved, even though the paper’s log-linear scaling curve says it should have. What’s the first thing you measure?”
Don’t say: “The base model isn’t strong enough, we need a bigger model.”
Why throwing more inference budget at a reasoning task quietly yields ninety thousand identical wrong answers, and how decomposing your flat curve into coverage vs. selection saves you from a six-fig.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.