LLM Inference Interview Questions, issue 4, Aug 2, 2026
The Best-of-N Paradox
Senior AI Engineer interview at Anthropic, and the interviewer asks:
“You’re reranking 16 agent trajectories with a critic model to boost your SWE-bench score. Accuracy went up. Now walk me through the exact shape of that accuracy-vs-rollouts curve, and tell me when this 16x spend is actually worth it in production.”
Don’t say: “More rollouts means more chances to get it right, so accuracy goes up.”
Why buying agent accuracy on a logarithmic curve silently destroys your inference budget, and the exact economic threshold where a 16x compute multiplier is actually justified.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.