RAG Interview Questions, issue 8, Jul 12, 2026

The IVF-PQ Compression Trick

Staff ML Engineer interview at Google, and the interviewer asks:

Your RAG system hits 99% recall on HNSW in the demo. Now it’s 200M vectors and your RAM bill just triggered a budget review. Do you keep HNSW? Why or why not?

Don’t say: HNSW is faster and more accurate, so we keep it and add more RAM.

Why blindly scaling HNSW is a scaling nightmare, and how to gracefully trade imperceptible coverage loss for massive infrastructure savings.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.