LLM Inference Interview Questions, issue 18, Aug 18, 2026
The Log-Linear Inference Trap
ML Engineer interview at Anthropic, and the interviewer asks:
“Best-of-16 rollouts with a critic reranker moves your agent from 20% to 32% on SWE-bench. Your PM wants it shipped. What do you tell him?”
Don’t say: “It’s a 60% relative improvement, let’s ship it, we just need to budget for the extra inference.”
Why using Best-of-N to boost agent performance quietly bankrupts your QPS budget, and the elite-level difference between buying benchmark points and shipping a viable AI product.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.