Advanced Reinforcement Learning Interview Questions, issue 9, Feb 4, 2026
The Perfect Classifier Trap
Senior RL Engineer interview at NVIDIA, and the interviewer asks:
“We trained a classifier to distinguish 𝘎𝘰𝘢𝘭 𝘙𝘦𝘢𝘤𝘩𝘦𝘥 vs. 𝘍𝘢𝘪𝘭𝘦𝘥 using 50 expert demos. It memorized the training set perfectly (100% Accuracy) in 10 epochs. But when we use this classifier as a reward signal, the robot learns absolutely nothing. Why?”
In RL, a reward model that memorizes expert demos creates a step-function objective that guarantees zero learning dynamics.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.