Computer Vision Interview Questions, issue 12, Jan 13, 2026

The Large Batch Generalization Trap

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

β€œWe just scaled our infrastructure to 4x our batch size (256 to 1024) to speed up training. We followed the π˜“π˜ͺ𝘯𝘦𝘒𝘳 𝘚𝘀𝘒𝘭π˜ͺ𝘯𝘨 π˜™π˜Άπ˜­π˜¦ and multiplied our π˜“π˜¦π˜’π˜³π˜―π˜ͺ𝘯𝘨 π˜™π˜’π˜΅π˜¦ by 4. But our test accuracy still degraded. What fundamental property of SGD did we accidentally kill?”

Why linear learning-rate scaling silently kills SGD’s implicit regularization and destroys test accuracy.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.