Advanced NLP Interview Questions, issue 23, Dec 28, 2025
The Curriculum Learning Trap
Staff Research Scientist interview at DeepSeek, and the interviewer asks:
โWe have three massive datasets: ๐๐ฆ๐ฏ๐ฆ๐ณ๐ข๐ญ ๐๐ฆ๐น๐ต, ๐๐ฐ๐ถ๐ณ๐ค๐ฆ ๐๐ฐ๐ฅ๐ฆ, and ๐ด๐ฑ๐ฆ๐ค๐ช๐ข๐ญ๐ช๐ป๐ฆ๐ฅ ๐๐ข๐ต๐ฉ ๐ฑ๐ณ๐ฐ๐ฃ๐ญ๐ฆ๐ฎ๐ด. To build a State-of-the-Art Math reasoner, in what order do you feed this data during pre-training, and why?โ
Donโt say: โJust shuffle them all together into one big dataset to avoid catastrophic forgetting.โ
Why shuffling General, Code, and Math data together silently caps reasoning performance and how staged pretraining unlocks true chain-of-thought.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.