Computer Vision Interview Questions, issue 7, Jan 8, 2026
The Receptive Field Trap
Senior Computer Vision Engineer interview at OpenAI, and the interviewer asks:
“In VGGNet, we replace a single 7x7 convolution with a stack of three 3x3 convolutions. Why?”
Why replacing a 7×7 convolution with three 3×3 layers isn’t about parameters — it’s about nonlinear expressivity.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.