Computer Vision Interview Questions, issue 6, Jan 7, 2026

The Model Capacity Trap

Senior MLE Engineer interview at Google DeepMind, and the interviewer asks:

Our new foundation model is overfitting severely on the training set. Should we cut the hidden dimension size from 4096 to 1024 to limit its capacity?

Why shrinking an overfitting network makes optimization harder, and why over-parameterization is the safer bet.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.