Computer Vision Interview Questions, issue 12, Jan 13, 2026
The Large Batch Generalization Trap
Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
βWe just scaled our infrastructure to 4x our batch size (256 to 1024) to speed up training. We followed the ππͺπ―π¦π’π³ ππ€π’ππͺπ―π¨ ππΆππ¦ and multiplied our ππ¦π’π³π―πͺπ―π¨ ππ’π΅π¦ by 4. But our test accuracy still degraded. What fundamental property of SGD did we accidentally kill?β
Why linear learning-rate scaling silently kills SGDβs implicit regularization and destroys test accuracy.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.