Issue 19, Oct 8, 2026

The Rejection Sampling Trap

Senior ML Engineer interview at Google DeepMind, and the interviewer asks:

“Your lead wants to skip RL entirely: sample from the model, keep the rollouts that pass the tests, fine-tune on them, repeat. When is that good enough to ship, and where does it hit a ceiling that policy gradient doesn’t?”

Don’t say: “Filtering by reward is basically RL, so it’ll work the same.”

Why filtering passed rollouts is secretly just policy gradient without a baseline and the hidden ceiling that halts your agent's progress.

Share this trapShare on LinkedIn

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.

More traps set at Google DeepMind