LLM System Design Interview, issue 69, Sep 13, 2026

The Data Plumbing Illusion

Senior Data Engineer interview at Anthropic, and the interviewer asks:

You swapped Common Crawl’s WET files for your own extraction off the raw WARC files. Same URL list, same filters, same model, but your benchmark scores moved. Your teammate says extraction is just plumbing. What do you tell him?

Don’t say: The extractor must have a bug, I’d diff the outputs.

Why treating HTML extraction as dumb plumbing silently purges your highest-signal reasoning tokens, and breaks every downstream quality filter in your pretraining pipeline.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.