Advanced NLP Interview Questions, issue 23, Dec 28, 2025

The Curriculum Learning Trap

Staff Research Scientist interview at DeepSeek, and the interviewer asks:

โ€œWe have three massive datasets: ๐˜Ž๐˜ฆ๐˜ฏ๐˜ฆ๐˜ณ๐˜ข๐˜ญ ๐˜›๐˜ฆ๐˜น๐˜ต, ๐˜š๐˜ฐ๐˜ถ๐˜ณ๐˜ค๐˜ฆ ๐˜Š๐˜ฐ๐˜ฅ๐˜ฆ, and ๐˜ด๐˜ฑ๐˜ฆ๐˜ค๐˜ช๐˜ข๐˜ญ๐˜ช๐˜ป๐˜ฆ๐˜ฅ ๐˜”๐˜ข๐˜ต๐˜ฉ ๐˜ฑ๐˜ณ๐˜ฐ๐˜ฃ๐˜ญ๐˜ฆ๐˜ฎ๐˜ด. To build a State-of-the-Art Math reasoner, in what order do you feed this data during pre-training, and why?โ€

Donโ€™t say: โ€œJust shuffle them all together into one big dataset to avoid catastrophic forgetting.โ€

Why shuffling General, Code, and Math data together silently caps reasoning performance and how staged pretraining unlocks true chain-of-thought.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.