LLM System Design Interview, issue 67, Sep 11, 2026

The Infinite Web Paradox

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

Your VP says: ‘We scraped the web in 2020 and got a great corpus. Just re-run the crawler and we’ll get the same data, only bigger.’ Why is that wrong, and how does it change your pre-training data strategy?

Don’t say: The web keeps growing, so a fresh crawl gives more tokens. We’ll just dedupe and filter harder.

More tokens online, less data you can actually touch: how consent walls and paywalls break naive scaling assumptions, and how to maximize the tokens you already have.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.