Advanced Reinforcement Learning Interview Questions, issue 3, Jan 29, 2026

The Covariate Shift Trap

Machine Learning Engineer interview at OpenAI, and the interviewer asks:

โ€œWe have a massive dataset of human expert demonstrations for this task. Why shouldnโ€™t we just stick with ๐˜๐˜ฎ๐˜ช๐˜ต๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ ๐˜“๐˜ฆ๐˜ข๐˜ณ๐˜ฏ๐˜ช๐˜ฏ๐˜จ (๐˜‰๐˜ฆ๐˜ฉ๐˜ข๐˜ท๐˜ช๐˜ฐ๐˜ณ ๐˜Š๐˜ญ๐˜ฐ๐˜ฏ๐˜ช๐˜ฏ๐˜จ)? Why take on the instability of ๐˜–๐˜ฏ๐˜ญ๐˜ช๐˜ฏ๐˜ฆ ๐˜—๐˜ฐ๐˜ญ๐˜ช๐˜ค๐˜บ ๐˜Ž๐˜ณ๐˜ข๐˜ฅ๐˜ช๐˜ฆ๐˜ฏ๐˜ต๐˜ด?โ€

Donโ€™t say: โ€œBecause Reinforcement Learning is just better...โ€

Imitation Learning collapses when the policy drifts off expert trajectories, while online policy gradients deliberately train on their own mistakes.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.