Advanced Reinforcement Learning Interview Questions, issue 8, Feb 3, 2026
The KL Regularization Trap
Senior AI Engineer interview at Anthropic, and the interviewer asks:
“Our Reward scores are climbing, but the 𝘒𝘓 𝘋𝘪𝘷𝘦𝘳𝘨𝘦𝘯𝘤𝘦 term is spiking. A junior engineer suggests setting the KL coefficient (Beta) to zero to unblock the model and maximize the reward faster. Do we approve the PR?”
When candidates treat KL as friction instead of a safety tether, they approve training loops that Goodhart themselves into gibberish.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.