Issue 17, Oct 6, 2026

The VLM Judge Trap

Senior AI Engineer interview at OpenAI, and the interviewer asks:

“Your computer-use agent runs real-world tasks where you can’t write programmatic verifiers. So you use a VLM judge with 85% human agreement as your RL reward. Is that safe? And what does the 15% disagreement turn into once the policy starts training against it?”

Don’t say: “85% is close to human-level. RL is robust to noisy rewards, so the errors average out.”

Why an 85% human agreement score quietly turns into a reward-hacking target, and how RL policies systematically exploit your evaluator's blind spots.

Share this trapShare on LinkedIn

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.

More traps set at OpenAI