Advanced NLP Interview Questions, issue 21, Dec 27, 2025

The PPO vs DPO Implementation Trap

Senior AI Engineer interview at OpenAI, and the interviewer asks:

โ€œOur engineers want to rip out ๐˜—๐˜—๐˜– (๐˜—๐˜ณ๐˜ฐ๐˜น๐˜ช๐˜ฎ๐˜ข๐˜ญ ๐˜—๐˜ฐ๐˜ญ๐˜ช๐˜ค๐˜บ ๐˜–๐˜ฑ๐˜ต๐˜ช๐˜ฎ๐˜ช๐˜ป๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ) and replace it with ๐˜‹๐˜—๐˜– (๐˜‹๐˜ช๐˜ณ๐˜ฆ๐˜ค๐˜ต ๐˜—๐˜ณ๐˜ฆ๐˜ง๐˜ฆ๐˜ณ๐˜ฆ๐˜ฏ๐˜ค๐˜ฆ ๐˜–๐˜ฑ๐˜ต๐˜ช๐˜ฎ๐˜ช๐˜ป๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ). They argue itโ€™s strictly better because it simplifies the stack. Do we approve the PR?โ€

Why replacing PPO with DPO is not a free lunch - and how gradient saturation turns โ€œsimplerโ€ into โ€œweaker.โ€

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.