Issue 17, Oct 6, 2026
The VLM Judge Trap
Senior AI Engineer interview at OpenAI, and the interviewer asks:
“Your computer-use agent runs real-world tasks where you can’t write programmatic verifiers. So you use a VLM judge with 85% human agreement as your RL reward. Is that safe? And what does the 15% disagreement turn into once the policy starts training against it?”
Don’t say: “85% is close to human-level. RL is robust to noisy rewards, so the errors average out.”
Why an 85% human agreement score quietly turns into a reward-hacking target, and how RL policies systematically exploit your evaluator's blind spots.
Share this trapShare on LinkedIn
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.