Issue 19, Oct 8, 2026
The Rejection Sampling Trap
Senior ML Engineer interview at Google DeepMind, and the interviewer asks:
“Your lead wants to skip RL entirely: sample from the model, keep the rollouts that pass the tests, fine-tune on them, repeat. When is that good enough to ship, and where does it hit a ceiling that policy gradient doesn’t?”
Don’t say: “Filtering by reward is basically RL, so it’ll work the same.”
Why filtering passed rollouts is secretly just policy gradient without a baseline and the hidden ceiling that halts your agent's progress.
Share this trapShare on LinkedIn
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.