Computer Vision Interview Questions, issue 1, Dec 31, 2025

The Translation Equivariance Efficiency Trap

Senior Computer Vision Engineer interview at Tesla, and the interviewer asks:

We all know 𝐂𝐍𝐍𝐬 are translation equivariant. But why exactly does that property make them exponentially more data-efficient than a 𝐅𝐮𝐥𝐥𝐲 𝐂𝐨𝐧𝐧𝐞𝐜𝐭𝐞𝐝 𝐧𝐞𝐭𝐰𝐨𝐫𝐤 for processing high-res images?

Why CNNs learn one visual feature once, while dense networks must relearn it at every pixel.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.