LLM System Design Interview, issue 19, Nov 16, 2025

Why โ€˜Train on the Internetโ€™ Guarantees a Trash Model

AI Engineer interview at OpenAI, and the interviewer asks:

โ€œA project plan budgets 1 day for data prep: ๐˜‹๐˜ฐ๐˜ธ๐˜ฏ๐˜ญ๐˜ฐ๐˜ข๐˜ฅ ๐˜Š๐˜ฐ๐˜ฎ๐˜ฎ๐˜ฐ๐˜ฏ ๐˜Š๐˜ณ๐˜ข๐˜ธ๐˜ญ. Why is this ๐˜ต๐˜ณ๐˜ข๐˜ช๐˜ฏ ๐˜ฐ๐˜ฏ ๐˜ต๐˜ฉ๐˜ฆ ๐˜ช๐˜ฏ๐˜ต๐˜ฆ๐˜ณ๐˜ฏ๐˜ฆ๐˜ต mindset a complete fantasy that guarantees a ๐ญ๐ซ๐š๐ฌ๐ก model ?โ€

Donโ€™t say: โ€œBecause the raw data is low-quality. You need to apply quality filters to remove spam, filter out harmful content, and deduplicate the data so the model doesnโ€™t memorize web pages.โ€

The hidden, multi-month pipeline - parsing, filtering, and deduping - that turns raw web trash into frontier-model intelligence.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.