Advanced Deep Learning Interview Questions, issue 3, Mar 24, 2026

The Leaderboard Overfitting Trap

Senior ML Engineer interview at Meta, and the interviewer asks:

Your team just ensembled 12 different deep learning models to squeeze out an extra 2% accuracy and secure the top spot on our internal leaderboard. Why is directly deploying this ‘winning’ submission a terrible idea for our live system, and what technique do you use instead?

Don’t say: It’s too computationally expensive to run 12 models, so it will cost the company too much money.

Winning offline metrics hides the fact that production systems penalize inference latency, orchestration overhead, and failure surface area.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.