LLM Inference Interview Questions, issue 19, Aug 19, 2026

The 10x Sampling Trap

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

You 10x’d your sampling budget on a hard reasoning task and the solve rate barely moved, even though the paper’s log-linear scaling curve says it should have. What’s the first thing you measure?

Don’t say: The base model isn’t strong enough, we need a bigger model.

Why throwing more inference budget at a reasoning task quietly yields ninety thousand identical wrong answers, and how decomposing your flat curve into coverage vs. selection saves you from a six-fig.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.