LLM System Design Interview, issue 19, Nov 16, 2025
Why โTrain on the Internetโ Guarantees a Trash Model
AI Engineer interview at OpenAI, and the interviewer asks:
โA project plan budgets 1 day for data prep: ๐๐ฐ๐ธ๐ฏ๐ญ๐ฐ๐ข๐ฅ ๐๐ฐ๐ฎ๐ฎ๐ฐ๐ฏ ๐๐ณ๐ข๐ธ๐ญ. Why is this ๐ต๐ณ๐ข๐ช๐ฏ ๐ฐ๐ฏ ๐ต๐ฉ๐ฆ ๐ช๐ฏ๐ต๐ฆ๐ณ๐ฏ๐ฆ๐ต mindset a complete fantasy that guarantees a ๐ญ๐ซ๐๐ฌ๐ก model ?โ
Donโt say: โBecause the raw data is low-quality. You need to apply quality filters to remove spam, filter out harmful content, and deduplicate the data so the model doesnโt memorize web pages.โ
The hidden, multi-month pipeline - parsing, filtering, and deduping - that turns raw web trash into frontier-model intelligence.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.