Advanced Reinforcement Learning Interview Questions, issue 9, Feb 4, 2026

The Perfect Classifier Trap

Senior RL Engineer interview at NVIDIA, and the interviewer asks:

We trained a classifier to distinguish 𝘎𝘰𝘢𝘭 𝘙𝘦𝘢𝘤𝘩𝘦𝘥 vs. 𝘍𝘢𝘪𝘭𝘦𝘥 using 50 expert demos. It memorized the training set perfectly (100% Accuracy) in 10 epochs. But when we use this classifier as a reward signal, the robot learns absolutely nothing. Why?

In RL, a reward model that memorizes expert demos creates a step-function objective that guarantees zero learning dynamics.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.