LLM System Design Interview, issue 74, Sep 18, 2026
The Verifier False Negative Trap
Senior AI Engineer interview at OpenAI, and the interviewer asks:
“Your RLVR math pipeline has a verifier, so reward hacking isn’t your problem. But accuracy has plateaued far below what manual review says the model can do. Where is ‘verifiable’ failing you?”
Don’t say: “The model isn’t learning the task, so I’d scale up RL compute or add harder problems.”
Everyone obsesses over the model cheating the reward, but the costliest RL failure is the reward cheating the model, silently penalizing correct reasoning right at your learning frontier.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.