LLM Inference Interview Questions, issue 18, Aug 18, 2026

The Log-Linear Inference Trap

ML Engineer interview at Anthropic, and the interviewer asks:

Best-of-16 rollouts with a critic reranker moves your agent from 20% to 32% on SWE-bench. Your PM wants it shipped. What do you tell him?

Don’t say: It’s a 60% relative improvement, let’s ship it, we just need to budget for the extra inference.

Why using Best-of-N to boost agent performance quietly bankrupts your QPS budget, and the elite-level difference between buying benchmark points and shipping a viable AI product.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.