LLM System Design Interview, issue 72, Sep 16, 2026
The Proxy Reward Trap
Senior ML Engineer interview at OpenAI, and the interviewer asks:
“Your PPO reward model scores climbed for six weeks while you kept adding compute. Human evals got worse. Why did more RL compute stop working, and what does your reward need before scaling RL actually pays off?”
Don’t say: “The learning rate was too high, so tune the KL coefficient and train longer.”
Why pouring GPU compute into PPO quietly degrades true output quality and how frontier labs catch reward overoptimization before burning millions on ghost gains.
The full answer, with the mechanism and the arithmetic, is free on Substack.