LLM Inference Interview Questions, issue 23, Aug 23, 2026
The Search-o1 Trick
Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
“Your agent retrieves mid-reasoning instead of once upfront. Recall@10 went up. Answer accuracy went down. Where’s the failure?”
Don’t say: “The retriever is pulling in irrelevant documents, so I’d tune the reranker.”
Why high recall quietly degrades agentic accuracy, and how shifting from raw document retrieval to mid-stream context compression saves your signal-to-noise ratio.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.