Advanced Reinforcement Learning Interview Questions, issue 3, Jan 29, 2026
The Covariate Shift Trap
Machine Learning Engineer interview at OpenAI, and the interviewer asks:
โWe have a massive dataset of human expert demonstrations for this task. Why shouldnโt we just stick with ๐๐ฎ๐ช๐ต๐ข๐ต๐ช๐ฐ๐ฏ ๐๐ฆ๐ข๐ณ๐ฏ๐ช๐ฏ๐จ (๐๐ฆ๐ฉ๐ข๐ท๐ช๐ฐ๐ณ ๐๐ญ๐ฐ๐ฏ๐ช๐ฏ๐จ)? Why take on the instability of ๐๐ฏ๐ญ๐ช๐ฏ๐ฆ ๐๐ฐ๐ญ๐ช๐ค๐บ ๐๐ณ๐ข๐ฅ๐ช๐ฆ๐ฏ๐ต๐ด?โ
Donโt say: โBecause Reinforcement Learning is just better...โ
Imitation Learning collapses when the policy drifts off expert trajectories, while online policy gradients deliberately train on their own mistakes.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.