Advanced Reinforcement Learning Interview Questions, issue 5, Jan 31, 2026
The Success-Only Dataset Trap
Research Scientist interview at Google DeepMind, and the interviewer asks:
โI have a dataset of reasoning traces, but theyโre all flawed. - ๐๐ณ๐ข๐ค๐ฆ ๐ ๐ด๐ต๐ข๐ณ๐ต๐ด ๐ธ๐ช๐ต๐ฉ ๐ฑ๐ฆ๐ณ๐ง๐ฆ๐ค๐ต ๐ญ๐ฐ๐จ๐ช๐ค ๐ฃ๐ถ๐ต ๐ฉ๐ข๐ญ๐ญ๐ถ๐ค๐ช๐ฏ๐ข๐ต๐ฆ๐ด ๐ต๐ฉ๐ฆ ๐ง๐ช๐ฏ๐ข๐ญ ๐ด๐ต๐ฆ๐ฑ (๐๐ข๐ช๐ญ). - ๐๐ณ๐ข๐ค๐ฆ ๐ ๐ด๐ต๐ข๐ณ๐ต๐ด ๐ธ๐ช๐ต๐ฉ ๐ข ๐ฎ๐ช๐ด๐ต๐ข๐ฌ๐ฆ ๐ฃ๐ถ๐ต ๐ญ๐ถ๐ค๐ฌ๐ช๐ญ๐บ ๐ณ๐ฆ๐ค๐ฐ๐ท๐ฆ๐ณ๐ด ๐ต๐ฐ ๐จ๐ฆ๐ต ๐ต๐ฉ๐ฆ ๐ณ๐ช๐จ๐ฉ๐ต ๐ข๐ฏ๐ด๐ธ๐ฆ๐ณ (๐๐ถ๐ค๐ค๐ฆ๐ด๐ด). Standard Imitation Learning (SFT) will ignore Trace A and clone Trace B, including its mistake. How do we train a model that outperforms both?โ
Filtering for correct answers discards optimal reasoning steps, preventing models from stitching together better policies than any single trace.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.