LLM Agents Interview Questions, issue 20, Mar 16, 2026

The Reward Signal Collapse Trap

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

Your RLHF pipeline relies on top-tier medical and legal experts to score outputs. But as the model scales, your PPO updates start degrading its reasoning accuracy rather than refining it. What is breaking down, and how do you fix it?

Don’t say: The PPO hyperparameters are unstable, or the reward model is overfitting. We need to add a stricter KL divergence penalty to keep the policy closer to the reference model.

When evaluators can't reliably judge advanced reasoning, PPO doesn't refine the model, it optimizes toward human-perceived correctness instead of actual correctness.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.