Advanced Reinforcement Learning Interview Questions, issue 4, Jan 30, 2026

The LLM-as-a-Judge Trap

Machine Learning Engineer interview at DeepSeek, and the interviewer asks:

β€œWe want to train a reasoning model using πƒπ’π«πžπœπ­ 𝐏𝐫𝐞𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝐎𝐩𝐭𝐒𝐦𝐒𝐳𝐚𝐭𝐒𝐨𝐧 (πƒππŽ), but we have zero budget for human annotators. How do we procedurally generate high-quality 𝘞π˜ͺ𝘯𝘯𝘦𝘳 𝘷𝘴. π˜“π˜°π˜΄π˜¦π˜³ pairs from the model’s own generations?”

Outsourcing preference labels to a stronger model just distills its bias, while execution-verified self-play produces cleaner gradients and scales without a teacher ceiling.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.