LLM Inference Interview Questions, issue 4, Aug 2, 2026

The Best-of-N Paradox

Senior AI Engineer interview at Anthropic, and the interviewer asks:

You’re reranking 16 agent trajectories with a critic model to boost your SWE-bench score. Accuracy went up. Now walk me through the exact shape of that accuracy-vs-rollouts curve, and tell me when this 16x spend is actually worth it in production.

Don’t say: More rollouts means more chances to get it right, so accuracy goes up.

Why buying agent accuracy on a logarithmic curve silently destroys your inference budget, and the exact economic threshold where a 16x compute multiplier is actually justified.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.