LLM Agents Interview Questions, issue 22, Mar 18, 2026

The Verifiable Reward Bypass Trap

Senior AI Engineer interview at OpenAI, and the interviewer asks:

You’re fine-tuning an LLM for instruction following (IFEval) using PPO. By step 400, your reward curve is steadily climbing, but your actual evaluation scores are tanking. How do you fix the reward pipeline without just training a massive 70B reward model?

Using a neural reward model for strict constraints is a category error, replace it with deterministic evaluation or guarantee systematic reward hacking.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.