Issue 13, Oct 1, 2026
The Functional Correctness Trap
Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
“Your RL reward for code generation is ‘all tests pass.’ Six weeks in, pass rates are up, but reviewers say the code is slow and full of unnecessary comments. What’s missing from your reward, and what breaks if you just add a readability judge?”
Don’t say: “Add a readability reward and a latency penalty.”
Why optimizing code models solely on unit tests quietly corrupts non-functional quality, and the gated reward strategy senior engineers use to prevent judge hacking.
Share this trapShare on LinkedIn
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.