RAG Interview Questions, issue 10, Jul 14, 2026

The Raw Vector Trap

Staff ML Engineer interview at Google, and the interviewer asks:

You and a teammate both ship IVF+PQ (Inverted File and Product Quantization) with identical memory budgets. Their recall is noticeably higher than yours. You quantized the raw vectors. They quantized the residual. Why does that one change beat you?

Don’t say: Residuals are smaller, so it compresses better.

Why quantizing raw embeddings silently tanks your retrieval recall, and how shifting your PQ budget to encode residuals buys you massive resolution at zero extra bytes.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.