LLM Inference Interview Questions, issue 23, Aug 23, 2026

The Search-o1 Trick

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

Your agent retrieves mid-reasoning instead of once upfront. Recall@10 went up. Answer accuracy went down. Where’s the failure?

Don’t say: The retriever is pulling in irrelevant documents, so I’d tune the reranker.

Why high recall quietly degrades agentic accuracy, and how shifting from raw document retrieval to mid-stream context compression saves your signal-to-noise ratio.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.