Generative Vision Interview Questions, issue 20, Jun 28, 2026

The Reward Hacking Trap

Machine Learning Engineer interview at OpenAI, and the interviewer asks:

You fine-tuned your image model against a reward model. Your alignment scores jumped 30%. But humans say the outputs got worse. What happened and how do you stop it?

Don’t say: The reward model must be undertrained, I’d collect more preference data.

Why chasing a 30% jump in alignment scores silently turns your image model into an adversarial example generator, and why the KL penalty is the only leash that keeps optimization honest.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.