Advanced Reinforcement Learning Interview Questions

Rewards, policies and RLHF. What breaks when the reward is a model.

25 traps, Jan 2026 to Feb 2026. Complete.

Set in interviews at Google DeepMind (7), OpenAI (7), Anthropic (4) and Boston Dynamics (1).

The panda for Advanced Reinforcement Learning Interview Questions, drawn at a desk

Each trap: the interviewer’s question, the answer most candidates give, and the mechanism that breaks it. The full answers are on Substack.

  1. Senior AI Engineer interview at Anthropic

    The Sparse Reward Trap

    You implemented Rejection Fine-Tuning (RFT) by sampling N solutions per problem and training on the correct ones. To push pass@1 accuracy, you drastically scale N, generating 100x more samples per prompt. Why does your test set error suddenly spike?

    Reinforcement Learning #25
    Feb 20, 2026

  2. Senior AI Robotics Engineer interview at Google DeepMind

    The Amortization Trap

    We swapped our Actor-Critic stack for pure Q-learning to simplify our architecture. In our 14-DoF continuous action space, why does the standard argmax(Q) operation completely shatter our 5ms inference latency budget, and how do you fix it?

    Reinforcement Learning #24
    Feb 19, 2026

  3. Senior Robotics Engineer interview at OpenAI

    The State Visitation Trap

    We want to minimize human interventions. So, we increase the episode length (H) from 1k to 100k steps in our SAC agent. We are collecting 100x more data per reset. Why does the policy’s success rate collapse to near-zero?

    Reinforcement Learning #23
    Feb 18, 2026

  4. Senior RLHF Engineer interview at OpenAI

    The Information Density Trap

    We have a $50k budget for human labeling. We need a reward model for ‘helpfulness.’ Do we pay humans to score responses on a 1-10 scale, or rank pairs (A > B)?

    Reinforcement Learning #22
    Feb 17, 2026

  5. Machine Learning Engineer interview at OpenAI

    The Happy Path Trap

    We are building an RL agent to grade student-coded video games (like Breakout). How do you design the reward function to catch the most bugs?

    Reinforcement Learning #21
    Feb 16, 2026

  6. Principal AI Engineer interview

    The Static CoT Trap

    We’re building a reasoning model like DeepSeek R1. We want the model to burn test-time compute exploring solutions for complex math, but answer instantly for ‘2+2’. How do you formulate the RL objective to achieve this adaptive behavior?

    Reinforcement Learning #20
    Feb 15, 2026

  7. Senior RL Engineer interview at OpenAI

    The Small-Batch Policy Gradient Trap

    We collected 6 robot trajectories. 5 failed (low reward). 1 succeeded (high reward). We run a vanilla Policy Gradient update on this small batch. What happens to the gradient?

    Reinforcement Learning #19
    Feb 14, 2026

  8. Senior AI Engineer interview at Google

    The Gumbel-Softmax Trap

    If the discrete token bottleneck is the main reason we need RLHF, why not just use the Gumbel-Softmax trick to make sampling differentiable and backpropagate end-to-end?

    Reinforcement Learning #18
    Feb 13, 2026

  9. Senior Engineer interview at OpenAI

    The Transitivity Assumption Trap

    How do you clean this data before training your Reward Model?

    Reinforcement Learning #17
    Feb 12, 2026

  10. Senior RL Engineer interview at OpenAI

    The Bootstrapping Bias Trap

    We accidentally initialized our Value Network to output -1000 for every state. We run one update step using Monte Carlo and one using Bootstrapping (TD-Learning). Which algorithm breaks immediately, and which one survives?

    Reinforcement Learning #16
    Feb 11, 2026

  11. Senior AI Research Engineer interview at Google DeepMind

    The Cold Start Exploration Trap

    We are training a new RL agent to manipulate a robot arm for a task like pouring water. A junior engineer suggests initializing with standard epsilon-greedy exploration to discover the first high-reward state. Why is this mathematically doomed, and what is the production-ready alternative?

    Reinforcement Learning #15
    Feb 10, 2026

  12. Senior RL Research Scientist interview at Google DeepMind

    The Local Randomness Trap

    We’re training an agent for a sparse-reward, long-horizon task. Your Epsilon-Greedy agent is flatlining and stuck in local optima. However, a Thompson Sampling agent solves it efficiently. Why? What is the fundamental difference in how they treat uncertainty?

    Reinforcement Learning #14
    Feb 9, 2026

  13. Senior Research Scientist interview at Google DeepMind

    The Dead Gradient Trap

    We’re training an end-to-end Meta-RL agent to find objects in a procedurally generated house. The loss curves are completely flat, the agent isn’t learning to explore or solve the task. Why is the gradient dead, and what is the fundamental coupling failure happening here?

    Reinforcement Learning #13
    Feb 8, 2026

  14. Senior RL Engineer interview at Google DeepMind

    The OOD Extrapolation Trap

    We have 50TB of static historical logs. If we run a standard off-policy algorithm (like Soft Actor-Critic) on this buffer without collecting new data, what happens to the Q-values?

    Reinforcement Learning #12
    Feb 7, 2026

  15. Principal ML Engineer interview at Anthropic

    The Context Leakage Trap

    We need to scale our RLHF dataset 10x. To maximize variety, we present human labelers with two model outputs generated from different user prompts (e.g., Response A to Prompt X vs. Response B to Prompt Y). We ask them to pick the better response.

    Reinforcement Learning #11
    Feb 6, 2026

  16. Senior RL Engineer interview at Google DeepMind

    The Boltzmann Collapse Trap

    You’re implementing Conservative Q-Learning (CQL). To penalize out-of-distribution actions, you need to find the actions with the highest Q-values. Should we spin up a separate optimizer network to hunt for these maximums?

    Reinforcement Learning #10
    Feb 5, 2026

  17. Senior RL Engineer interview at NVIDIA

    The Perfect Classifier Trap

    We trained a classifier to distinguish 𝘎𝘰𝘢𝘭 𝘙𝘦𝘢𝘤𝘩𝘦𝘥 vs. 𝘍𝘢𝘪𝘭𝘦𝘥 using 50 expert demos. It memorized the training set perfectly (100% Accuracy) in 10 epochs. But when we use this classifier as a reward signal, the robot learns absolutely nothing. Why?

    Reinforcement Learning #9
    Feb 4, 2026

  18. Senior AI Engineer interview at Anthropic

    The KL Regularization Trap

    Our Reward scores are climbing, but the 𝘒𝘓 𝘋𝘪𝘷𝘦𝘳𝘨𝘦𝘯𝘤𝘦 term is spiking. A junior engineer suggests setting the KL coefficient (Beta) to zero to unblock the model and maximize the reward faster. Do we approve the PR?

    Reinforcement Learning #8
    Feb 3, 2026

  19. Senior AI Engineer interview at Boston Dynamics

    The Dynamics Invariance Trap

    We have 10k hours of data from a robot walking on concrete. We want to use Hindsight Relabeling (HER) to bootstrap a new policy for walking on sand. Is this a good idea?

    Reinforcement Learning #7
    Feb 2, 2026

  20. Senior AI Engineer interview at NVIDIA Robotics

    The Initialization Gap Trap

    We trained Policy A (Boil Water) to 99% accuracy. We trained Policy B (Find Pasta) to 99% accuracy. Both work perfectly in isolation. But when we run them in sequence (A → B), the robot fails immediately. Why?

    Reinforcement Learning #6
    Feb 1, 2026

  21. Research Scientist interview at Google DeepMind

    The Success-Only Dataset Trap

    I have a dataset of reasoning traces, but they’re all flawed. - 𝘛𝘳𝘢𝘤𝘦 𝘈 𝘴𝘵𝘢𝘳𝘵𝘴 𝘸𝘪𝘵𝘩 𝘱𝘦𝘳𝘧𝘦𝘤𝘵 𝘭𝘰𝘨𝘪𝘤 𝘣𝘶𝘵 𝘩𝘢𝘭𝘭𝘶𝘤𝘪𝘯𝘢𝘵𝘦𝘴 𝘵𝘩𝘦 𝘧𝘪𝘯𝘢𝘭 𝘴𝘵𝘦𝘱 (𝘍𝘢𝘪𝘭). - 𝘛𝘳𝘢𝘤𝘦 𝘉 𝘴𝘵𝘢𝘳𝘵𝘴 𝘸𝘪𝘵𝘩 𝘢 𝘮𝘪𝘴𝘵𝘢𝘬𝘦 𝘣𝘶𝘵 𝘭𝘶𝘤𝘬𝘪𝘭𝘺 𝘳𝘦𝘤𝘰𝘷𝘦𝘳𝘴 𝘵𝘰 𝘨𝘦𝘵 𝘵𝘩𝘦 𝘳𝘪𝘨𝘩𝘵 𝘢𝘯𝘴𝘸𝘦𝘳 (𝘚𝘶𝘤𝘤𝘦𝘴𝘴). Standard Imitation Learning (SFT) will ignore Trace A and clone Trace B, including its mistake. How do we train a model that outperforms both?

    Reinforcement Learning #5
    Jan 31, 2026

  22. Machine Learning Engineer interview at DeepSeek

    The LLM-as-a-Judge Trap

    We want to train a reasoning model using 𝐃𝐢𝐫𝐞𝐜𝐭 𝐏𝐫𝐞𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝐎𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧 (𝐃𝐏𝐎), but we have zero budget for human annotators. How do we procedurally generate high-quality 𝘞𝘪𝘯𝘯𝘦𝘳 𝘷𝘴. 𝘓𝘰𝘴𝘦𝘳 pairs from the model’s own generations?

    Reinforcement Learning #4
    Jan 30, 2026

  23. Machine Learning Engineer interview at OpenAI

    The Covariate Shift Trap

    We have a massive dataset of human expert demonstrations for this task. Why shouldn’t we just stick with 𝘐𝘮𝘪𝘵𝘢𝘵𝘪𝘰𝘯 𝘓𝘦𝘢𝘳𝘯𝘪𝘯𝘨 (𝘉𝘦𝘩𝘢𝘷𝘪𝘰𝘳 𝘊𝘭𝘰𝘯𝘪𝘯𝘨)? Why take on the instability of 𝘖𝘯𝘭𝘪𝘯𝘦 𝘗𝘰𝘭𝘪𝘤𝘺 𝘎𝘳𝘢𝘥𝘪𝘦𝘯𝘵𝘴?

    Reinforcement Learning #3
    Jan 29, 2026

  24. Machine Learning Engineer interview at Tesla

    The Mean Collapse Trap

    We have an imitation learning agent that is underfitting complex human driving data. A junior engineer suggests scaling the backbone network size by 10x to 𝘪𝘯𝘤𝘳𝘦𝘢𝘴𝘦 𝘤𝘢𝘱𝘢𝘤𝘪𝘵𝘺. We are currently using a simple Gaussian output head. Why will scaling the network fail to solve the problem, no matter how much compute you throw at it?

    Reinforcement Learning #2
    Jan 28, 2026

  25. Machine Learning Engineer interview at Anthropic

    The Stationarity Trap

    In Supervised Learning, we assume data is IID (Independent and Identically Distributed). Why does applying this assumption to a Reinforcement Learning agent, like a coding assistant, cause catastrophic failure?

    Reinforcement Learning #1
    Jan 27, 2026

Get the next one

Free on Substack. Unsubscribe in one click.