LLM System Design Interview, issue 8, Nov 6, 2025
The Contaminated Benchmark Trap
Lead AI Engineer interview at Anthropic, and the interviewer asks:
βOur new model just hit 95% on MMLU, beating GPT-4. The marketing team is drafting a press release. As the engineering lead, whatβs the π§πͺπ³π΄π΅ π΅π©πͺπ―π¨ you check for that could invalidate this result?β
Donβt say: βIβd check for train-test overlap.β
When 95% on MMLU doesnβt mean youβve built a smarter model - it means your training data leaked the exam answers. How to detect semantic contamination before your press release backfires.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.