Advanced NLP Interview Questions

Tokenizers, attention, fine-tuning and evaluation.

25 traps, Dec 2025. Complete.

Set in interviews at Google DeepMind (8), OpenAI (7), Meta (4) and Google (2).

The cat for Advanced NLP Interview Questions, drawn at a desk

Each trap: the interviewerโ€™s question, the answer most candidates give, and the mechanism that breaks it. The full answers are on Substack.

  1. Senior NLP Engineer interview at Google DeepMind

    The Back-Translation Direction Trap

    โ€œWe need to improve our ๐˜‘๐˜ข๐˜ฑ๐˜ข๐˜ฏ๐˜ฆ๐˜ด๐˜ฆ-๐˜ต๐˜ฐ-๐˜Œ๐˜ฏ๐˜จ๐˜ญ๐˜ช๐˜ด๐˜ฉ translation model. We have 10k parallel pairs and 1 billion lines of monolingual English text. To use ๐๐š๐œ๐ค-๐“๐ซ๐š๐ง๐ฌ๐ฅ๐š๐ญ๐ข๐จ๐ง effectively, which direction do we generate data, and exactly how do we pair it for training?โ€

    NLP #25
    Dec 30, 2025

  2. Senior AI Engineer interview at Anthropic

    The Confidence Calibration Trap

    โ€œWeโ€™re bleeding money on inference. We want to build a ๐Œ๐จ๐๐ž๐ฅ ๐‚๐š๐ฌ๐œ๐š๐๐ž (๐…๐ซ๐ฎ๐ ๐š๐ฅ๐†๐๐“) system, route easy queries to Llama-7B, and only send the hard stuff to GPT-4. What is the actual engineering bottleneck that makes this unreliable in production?โ€

    NLP #24
    Dec 29, 2025

  3. Staff Research Scientist interview at DeepSeek

    The Curriculum Learning Trap

    โ€œWe have three massive datasets: ๐˜Ž๐˜ฆ๐˜ฏ๐˜ฆ๐˜ณ๐˜ข๐˜ญ ๐˜›๐˜ฆ๐˜น๐˜ต, ๐˜š๐˜ฐ๐˜ถ๐˜ณ๐˜ค๐˜ฆ ๐˜Š๐˜ฐ๐˜ฅ๐˜ฆ, and ๐˜ด๐˜ฑ๐˜ฆ๐˜ค๐˜ช๐˜ข๐˜ญ๐˜ช๐˜ป๐˜ฆ๐˜ฅ ๐˜”๐˜ข๐˜ต๐˜ฉ ๐˜ฑ๐˜ณ๐˜ฐ๐˜ฃ๐˜ญ๐˜ฆ๐˜ฎ๐˜ด. To build a State-of-the-Art Math reasoner, in what order do you feed this data during pre-training, and why?โ€

    NLP #23
    Dec 28, 2025

  4. Senior ML Engineer interview at Google

    The Inter-Annotator Agreement Trap

    โ€œWeโ€™re building a toxicity detection dataset where only 1% of comments are actually toxic. We hired two annotators. Their inter-annotator agreement is 99%. Are we good to go?โ€

    NLP #22
    Dec 27, 2025

  5. Senior AI Engineer interview at OpenAI

    The PPO vs DPO Implementation Trap

    โ€œOur engineers want to rip out ๐˜—๐˜—๐˜– (๐˜—๐˜ณ๐˜ฐ๐˜น๐˜ช๐˜ฎ๐˜ข๐˜ญ ๐˜—๐˜ฐ๐˜ญ๐˜ช๐˜ค๐˜บ ๐˜–๐˜ฑ๐˜ต๐˜ช๐˜ฎ๐˜ช๐˜ป๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ) and replace it with ๐˜‹๐˜—๐˜– (๐˜‹๐˜ช๐˜ณ๐˜ฆ๐˜ค๐˜ต ๐˜—๐˜ณ๐˜ฆ๐˜ง๐˜ฆ๐˜ณ๐˜ฆ๐˜ฏ๐˜ค๐˜ฆ ๐˜–๐˜ฑ๐˜ต๐˜ช๐˜ฎ๐˜ช๐˜ป๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ). They argue itโ€™s strictly better because it simplifies the stack. Do we approve the PR?โ€

    NLP #21
    Dec 27, 2025

  6. Senior AI Engineer interview at Meta

    The Quantization Gradient Trap

    โ€œWe need to switch to ๐๐ฎ๐š๐ง๐ญ๐ข๐ณ๐š๐ญ๐ข๐จ๐ง ๐€๐ฐ๐š๐ซ๐ž ๐“๐ซ๐š๐ข๐ง๐ข๐ง๐  (๐๐€๐“) because post-training quantization is tanking our accuracy. But the rounding operation (Float -> Int8) is a step function with a derivative of zero. How do you actually backpropagate gradients through it to update the weights?โ€

    NLP #20
    Dec 26, 2025

  7. Senior AI Engineer interview at NVIDIA

    The QLoRA Compute Tax Trap

    โ€œWe switched from standard FP16 fine-tuning to QLoRA (4-bit quantization) to save memory. The model fits now, but training speed hasnโ€™t improved, itโ€™s actually slightly slower. Why didnโ€™t reducing precision by 4x result in a 4x speedup?โ€

    NLP #19
    Dec 25, 2025

  8. Senior AI Engineer interview at Google

    The Perplexity Tokenizer Trap

    โ€œWe ran an eval on a fixed dataset. Llama3 achieved a perplexity of 2.1, while Gemma3 scored 2.4. Which model is the better probability estimator, and which one do we deploy?โ€

    NLP #18
    Dec 24, 2025

  9. Senior ML Engineer interview at OpenAI

    The Sparse Gradient Trap

    โ€œWeโ€™re training a model on a massive vocabulary. Some critical domain terms appear only once every 10,000 documents. Why will standard SGD fail to learn weights for these rare features, and how does Adam specifically fix this?โ€

    NLP #17
    Dec 23, 2025

  10. Senior AI Engineer interview at Google DeepMind

    The Hinge Loss Confidence Trap

    โ€œWeโ€™re training a massive binary text classifier. A junior engineer suggests using Hinge Loss because it creates a ๐˜ฎ๐˜ข๐˜น ๐˜ฎ๐˜ข๐˜ณ๐˜จ๐˜ช๐˜ฏ and stops updating once a sample is correct, theoretically improving training stability. Why do we still prefer ๐’๐ข๐ ๐ฆ๐จ๐ข๐ + ๐‹๐จ๐  ๐‹๐ข๐ค๐ž๐ฅ๐ข๐ก๐จ๐จ๐ in production, specifically regarding the gradient signal on ๐˜ค๐˜ฐ๐˜ณ๐˜ณ๐˜ฆ๐˜ค๐˜ต examples?โ€

    NLP #16
    Dec 22, 2025

  11. Senior AI Engineer interview at Meta

    The Positional Encoding Wall

    โ€œWe used ๐˜™๐˜ฐ๐˜—๐˜Œ (๐˜™๐˜ฐ๐˜ต๐˜ข๐˜ณ๐˜บ ๐˜—๐˜ฐ๐˜ด๐˜ช๐˜ต๐˜ช๐˜ฐ๐˜ฏ๐˜ข๐˜ญ ๐˜Œ๐˜ฎ๐˜ฃ๐˜ฆ๐˜ฅ๐˜ฅ๐˜ช๐˜ฏ๐˜จ๐˜ด) for Llama instead of standard absolute learned embeddings. Apart from the math, what is the critical advantage RoPE offers when we need to run inference on sequences longer than what we trained on?โ€

    NLP #15
    Dec 21, 2025

  12. AI Engineer interview at Meta

    The Tokenization Brittleness Trap

    โ€œWe deployed a Llama-3 based app. We removed a single whitespace in the prompt template, and our benchmark accuracy tanked by 12%. Why is the model so brittle to a simple format change, and why didnโ€™t instruction tuning prevent this?โ€

    NLP #14
    Dec 20, 2025

  13. Senior AI Engineer interview at OpenAI

    The Knowledge Distillation Trap

    โ€œWe need to distill a massive 10-model ensemble into a single small model for low-latency serving. Why is training the student on the ensembleโ€™s final output tokens a complete waste of compute?โ€

    NLP #13
    Dec 19, 2025

  14. Senior Machine Learning Engineer interview at Google DeepMind

    The Optimizer State Memory Trap

    โ€œYou just switched a 7B parameter training run from SGD to Adam to speed up convergence. The model size is identical, but the cluster immediately crashes with a ๐˜Š๐˜œ๐˜‹๐˜ˆ ๐˜–๐˜ถ๐˜ต-๐˜–๐˜ง-๐˜”๐˜ฆ๐˜ฎ๐˜ฐ๐˜ณ๐˜บ (๐˜–๐˜–๐˜”) error. Why?โ€

    NLP #12
    Dec 18, 2025

  15. Senoir Machine Learning Engineer interview at Google DeepMind

    The Argmax Deadlock Trap

    โ€œWe are building a massive Mixture of Experts (MoE) model. To maximize training throughput on our H100 clusters, we want to route each token to only the single best expert (k=1). Is this a valid strategy?โ€

    NLP #11
    Dec 17, 2025

  16. Machine Learning Engineer interview at OpenAI

    The Contrastive Batch Size Trap

    โ€œWe need to train a specialized CLIP model for medical imaging from scratch. You have a node of 8 A100s. What batch size do you configure?โ€

    NLP #10
    Dec 16, 2025

  17. Senior NLP Engineer interview at Google DeepMind

    The Tokenization Trap in Semitic Languages

    โ€œWe need to adapt our English-centric LLM to support Arabic and Hebrew. How do you adjust the tokenizer?โ€

    NLP #9
    Dec 15, 2025

  18. Senior Machine Learning Engineer interview at OpenAI

    The Dropout Inference Trap

    โ€œWe implemented a custom Dropout layer from scratch. How do you handle it during inference?โ€

    NLP #8
    Dec 14, 2025

  19. Senior ML Engineer interview at Meta

    The Exploding Gradient Trap

    โ€œYouโ€™re training a 7B parameter Llama-style model. In the first 1000 steps, your gradients start oscillating wildly and the loss spikes. How do you fix it?โ€

    NLP #7
    Dec 13, 2025

  20. Senior ML Engineer interview at Google DeepMind

    The LoRA Initialization Trap

    โ€œYou are implementing ๐‹๐จ๐‘๐€ (๐‹๐จ๐ฐ-๐‘๐š๐ง๐ค ๐€๐๐š๐ฉ๐ญ๐š๐ญ๐ข๐จ๐ง) from scratch. How do you initialize the down-projection matrix A and the up-projection matrix B?โ€

    NLP #6
    Dec 13, 2025

  21. Machine Learning Engineer interview at OpenAI

    The Speculative Decoding Illusion

    โ€œWe need to optimize inference for batch size 128. Should we use Speculative Decoding?โ€

    NLP #5
    Dec 12, 2025

  22. final round ML Engineer interview at Google DeepMind

    The WEAT Bias Detection Trap

    โ€œHow do you prove your word embeddings arenโ€™t biased before we ship?โ€

    NLP #4
    Dec 11, 2025

  23. Senior ML Engineer interview at OpenAI

    The Attention Entropy Illusion

    โ€œHow do we use the Attention mechanismโ€™s weights to measure the modelโ€™s uncertainty?โ€

    NLP #3
    Dec 11, 2025

  24. Senior ML Engineer interview at NVIDIA

    The Gradient Shockwave Trap

    โ€œYou attach a new, random linear head to a pre-trained Transformer. Do you unfreeze all layers and start backprop immediately?โ€

    NLP #2
    Dec 9, 2025

  25. Senior ML Engineer interview at Google DeepMind

    The Learning Rate Warm-Up Trap

    โ€œWe are training a ๐˜›๐˜ณ๐˜ข๐˜ฏ๐˜ด๐˜ง๐˜ฐ๐˜ณ๐˜ฎ๐˜ฆ๐˜ณ from scratch using ๐˜ˆ๐˜ฅ๐˜ข๐˜ฎ. We set a constant Learning Rate of 1e-3. Predict the first 1000 steps.โ€

    NLP #1
    Dec 9, 2025

Get the next one

Free on Substack. Unsubscribe in one click.