Computer Vision Interview Questions, issue 7, Jan 8, 2026

The Receptive Field Trap

Senior Computer Vision Engineer interview at OpenAI, and the interviewer asks:

In VGGNet, we replace a single 7x7 convolution with a stack of three 3x3 convolutions. Why?

Why replacing a 7×7 convolution with three 3×3 layers isn’t about parameters — it’s about nonlinear expressivity.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.