Advanced Reinforcement Learning Interview Questions, issue 5, Jan 31, 2026

The Success-Only Dataset Trap

Research Scientist interview at Google DeepMind, and the interviewer asks:

โ€œI have a dataset of reasoning traces, but theyโ€™re all flawed. - ๐˜›๐˜ณ๐˜ข๐˜ค๐˜ฆ ๐˜ˆ ๐˜ด๐˜ต๐˜ข๐˜ณ๐˜ต๐˜ด ๐˜ธ๐˜ช๐˜ต๐˜ฉ ๐˜ฑ๐˜ฆ๐˜ณ๐˜ง๐˜ฆ๐˜ค๐˜ต ๐˜ญ๐˜ฐ๐˜จ๐˜ช๐˜ค ๐˜ฃ๐˜ถ๐˜ต ๐˜ฉ๐˜ข๐˜ญ๐˜ญ๐˜ถ๐˜ค๐˜ช๐˜ฏ๐˜ข๐˜ต๐˜ฆ๐˜ด ๐˜ต๐˜ฉ๐˜ฆ ๐˜ง๐˜ช๐˜ฏ๐˜ข๐˜ญ ๐˜ด๐˜ต๐˜ฆ๐˜ฑ (๐˜๐˜ข๐˜ช๐˜ญ). - ๐˜›๐˜ณ๐˜ข๐˜ค๐˜ฆ ๐˜‰ ๐˜ด๐˜ต๐˜ข๐˜ณ๐˜ต๐˜ด ๐˜ธ๐˜ช๐˜ต๐˜ฉ ๐˜ข ๐˜ฎ๐˜ช๐˜ด๐˜ต๐˜ข๐˜ฌ๐˜ฆ ๐˜ฃ๐˜ถ๐˜ต ๐˜ญ๐˜ถ๐˜ค๐˜ฌ๐˜ช๐˜ญ๐˜บ ๐˜ณ๐˜ฆ๐˜ค๐˜ฐ๐˜ท๐˜ฆ๐˜ณ๐˜ด ๐˜ต๐˜ฐ ๐˜จ๐˜ฆ๐˜ต ๐˜ต๐˜ฉ๐˜ฆ ๐˜ณ๐˜ช๐˜จ๐˜ฉ๐˜ต ๐˜ข๐˜ฏ๐˜ด๐˜ธ๐˜ฆ๐˜ณ (๐˜š๐˜ถ๐˜ค๐˜ค๐˜ฆ๐˜ด๐˜ด). Standard Imitation Learning (SFT) will ignore Trace A and clone Trace B, including its mistake. How do we train a model that outperforms both?โ€

Filtering for correct answers discards optimal reasoning steps, preventing models from stitching together better policies than any single trace.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.