Issue 11, Sep 30, 2026
The Near-Duplicate Leakage Trap
Senior ML Engineer interview at Google DeepMind, and the interviewer asks:
“You scraped GitHub, held out a random 5% of files as your test set, and your new code model just posted a big jump on it. Before you announce anything, what’s the first thing you’d suspect about how that split was built?”
Don’t say: “Overfitting. I’d check the training curves and add regularization.”
Why vendored libraries and fork networks silently inflate your evaluation metrics, and how to structure group-aware, temporal splits that measure true generalization.
Share this trapShare on LinkedIn
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.