RAG Interview Questions

Chunking, indexing, ranking and evals. Where retrieval quietly loses the answer.

25 traps, Jul 2026. Complete.

Set in interviews at Google (12), Anthropic (7), Pinterest (2) and Google DeepMind (1).

Each trap: the interviewer’s question, the answer most candidates give, and the mechanism that breaks it. The full answers are on Substack.

  1. Senior ML Engineer interview at Anthropic

    The "Lost in the Middle" Trap

    Your retriever pulls top-10 chunks, reranks them, and you feed all 10 to the LLM to be safe. Walk me through exactly what breaks, and when you’d retrieve fewer documents on purpose.

    RAG #25
    Jul 29, 2026

  2. Senior ML Engineer interview at Microsoft

    The Resolution Limit Trap

    You built GraphRAG with Leiden for community detection. It’s deterministic, gives clean hierarchies, everyone’s happy. Now your retrieval quality is degrading on certain queries. When does maximizing modularity actively hurt you?

    RAG #24
    Jul 28, 2026

  3. Senior ML Engineer interview at Anthropic

    The GraphRAG Evaluation Trap

    Your GraphRAG system beat vanilla RAG on summarization benchmarks. So why should I trust it on multi-hop reasoning?

    RAG #23
    Jul 27, 2026

  4. ML Engineer interview at Google DeepMind

    The Untyped Edge Paradox

    Your GraphRAG system uses a generic related_to edge to catch messy connections. Recall jumped. Six months later, retrieval quality is quietly rotting. What went wrong?

    RAG #22
    Jul 26, 2026

  5. Senior ML Engineer interview at Google

    The GraphRAG Scaling Trap

    Your GraphRAG pipeline uses an LLM to extract entity-relationship triplets. It’s flawless in the demo. Now you’re ingesting 50K docs a week and your indexing bill is on fire. What did you trade away when you picked LLM extraction?

    RAG #21
    Jul 25, 2026

  6. Staff ML Engineer interview at Google

    The Feedback Loop Trap

    Your REALM-style retriever learns from its own retrieval signal, no human labels. What breaks after six months in production?

    RAG #20
    Jul 24, 2026

  7. Senior ML Engineer interview at Google

    The Concatenation Trap

    Your conversational RAG has five query rewrite strategies. Why not just run all five and concatenate the outputs?

    RAG #19
    Jul 23, 2026

  8. Senior ML Engineer interview at Google

    The Hop-Count Paradox

    Adaptive RAG routes queries by predicted hop count, but nobody labels this question needs 2 hops. The paper bootstraps silver labels by brute-forcing every hop count and keeping whichever one got the right answer. What’s the ceiling that creates?

    RAG #18
    Jul 22, 2026

  9. Staff ML Engineer interview at OpenAI

    The Lexical Replacement Trap

    Your team wants to swap lexical query expansion for semantic expansion to boost precision. Walk me through why that’s the wrong way to do the migration.

    RAG #17
    Jul 21, 2026

  10. Staff ML Engineer interview at Anthropic

    The Green Dashboard Paradox

    Your Corrective-RAG system has an AI evaluator scoring its own retrievals. Six months in, answer quality is degrading, but your dashboards are green. What’s happening, and how would you have prevented it on day one?

    RAG #16
    Jul 20, 2026

  11. Senior ML Engineer interview at Anthropic

    The Ambiguity Relocation Trap

    You added an LLM query rewriter to fix ambiguous queries. Recall went up in your offline eval. So why did production accuracy quietly drop three weeks later?

    RAG #15
    Jul 19, 2026

  12. Staff ML Engineer interview at Anthropic

    The Iterative Retrieval Trap

    Your iterative RAG boosted multi-hop accuracy, but p99 latency tripled and inference costs are bleeding. A PM suggests capping retrieval at 3 loops. Why is that the wrong fix, and what should actually decide when to stop?

    RAG #14
    Jul 18, 2026

  13. Senior AI Engineer interview at Google

    The Flat Index Trap

    Your RAG pipeline scores 90% on factual lookups but collapses on ‘summarize this 300-page report.’ Why, and don’t tell me it’s chunk size?

    RAG #13
    Jul 17, 2026

  14. Staff ML Engineer interview at Pinterest

    The LSH Recall Paradox

    Your LSH index is missing 30% of true neighbors. You doubled the number of hash tables and recall barely moved. What’s actually happening?

    RAG #12
    Jul 16, 2026

  15. Senior AI Engineer interview at Meta

    The Zombie Hub Trap

    Your HNSW index takes a delete every few seconds in production. Do you remove the node and repair the graph?

    RAG #11
    Jul 15, 2026

  16. Staff ML Engineer interview at Google

    The Raw Vector Trap

    You and a teammate both ship IVF+PQ (Inverted File and Product Quantization) with identical memory budgets. Their recall is noticeably higher than yours. You quantized the raw vectors. They quantized the residual. Why does that one change beat you?

    RAG #10
    Jul 14, 2026

  17. Senior ML Engineer interview at Pinterest

    The Equal Split Trap

    You’re building a Product Quantization index. You split each vector into M subvectors and run k-means on each. A teammate says the split point doesn’t matter because the information is spread uniformly. Do you ship it?

    RAG #9
    Jul 13, 2026

  18. Staff ML Engineer interview at Google

    The IVF-PQ Compression Trick

    Your RAG system hits 99% recall on HNSW in the demo. Now it’s 200M vectors and your RAM bill just triggered a budget review. Do you keep HNSW? Why or why not?

    RAG #8
    Jul 12, 2026

  19. Senior AI Engineer interview at Google

    The Prestige Trap

    Your RAG system pulls from a SQL database, a private company corpus, and web search. The answers now contradict each other. How do you fix source selection?

    RAG #7
    Jul 11, 2026

  20. Senior ML Engineer interview at Anthropic

    The Reformulation Trap

    Your RAG system passes the raw user query straight to the retriever. Recall looks fine in your evals, but it collapses on real traffic, ambiguous and multi-hop questions especially. Before you touch the retriever, what does your query layer actually need to decide?

    RAG #6
    Jul 10, 2026

  21. Search Infrastructure Engineer interview at Google

    The Leaderboard Trap

    We benchmarked dense retrieval against BM25 on our eval set. Dense won on every metric. Do we rip out sparse retrieval entirely?

    RAG #5
    Jul 9, 2026

  22. Staff ML Engineer interview at Google

    The Parametric Memory Trap

    We want to swap our BM25 + dense pipeline for generative retrieval. Our corpus gains 50,000 documents every Monday. Why might that be a dealbreaker?

    RAG #4
    Jul 8, 2026

  23. Senior ML Engineer interview at Anthropic

    The Multi-Vector Trap

    You shipped ColBERT-style multi-vector retrieval because it won on nDCG@10. Two weeks later, p99 latency tripled and your index storage 30x’d. When is multi-vector actually worth that, and what exactly did you lose when you collapsed passages into single vectors before?

    RAG #3
    Jul 7, 2026

  24. Staff ML Engineer interview at Google

    The Semantic Chunking Trap

    Your RAG benchmark shows semantic chunking beating fixed-size by 8 points. You ship it. Production retrieval quality doesn’t move an inch. What happened?

    RAG #2
    Jul 6, 2026

  25. Senior ML Engineer interview at Google

    The Averaging Trap

    You’ve got three retrievers, BM25, a dense embedding model, and a rerank pass, and their relevance scores live on completely different scales. How do you merge them into one ranked list?

    RAG #1
    Jul 5, 2026

Get the next one

Free on Substack. Unsubscribe in one click.