Computer Vision Interview Questions, issue 1, Dec 31, 2025
The Translation Equivariance Efficiency Trap
Senior Computer Vision Engineer interview at Tesla, and the interviewer asks:
“We all know 𝐂𝐍𝐍𝐬 are translation equivariant. But why exactly does that property make them exponentially more data-efficient than a 𝐅𝐮𝐥𝐥𝐲 𝐂𝐨𝐧𝐧𝐞𝐜𝐭𝐞𝐝 𝐧𝐞𝐭𝐰𝐨𝐫𝐤 for processing high-res images?”
Why CNNs learn one visual feature once, while dense networks must relearn it at every pixel.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.