RAG Interview Questions, issue 2, Jul 6, 2026

The Semantic Chunking Trap

Staff ML Engineer interview at Google, and the interviewer asks:

Your RAG benchmark shows semantic chunking beating fixed-size by 8 points. You ship it. Production retrieval quality doesn’t move an inch. What happened?

Don’t say: Must be a bug in the embedding pipeline

Why synthetic benchmarks trick you into shipping expensive text splitters, and how understanding topic coherence saves your compute budget without losing a single point of retrieval quality.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.