RAG Interview Questions, issue 2, Jul 6, 2026
The Semantic Chunking Trap
Staff ML Engineer interview at Google, and the interviewer asks:
“Your RAG benchmark shows semantic chunking beating fixed-size by 8 points. You ship it. Production retrieval quality doesn’t move an inch. What happened?”
Don’t say: “Must be a bug in the embedding pipeline”
Why synthetic benchmarks trick you into shipping expensive text splitters, and how understanding topic coherence saves your compute budget without losing a single point of retrieval quality.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.