Generative Vision Interview Questions, issue 10, Jun 18, 2026
The Receptive Field Illusion
Senior ML Engineer interview at Midjourney, and the interviewer asks:
“Your teammate wants to ship a U-Net for your new high-res image model because ‘convolutions capture both local and global features.’ You disagree. Defend it.”
Don’t say: “Transformers are just better, everyone uses DiT now.”
Why stacking convolutions to build global context quietly turns distant features into mush, and how trading inductive bias for self-attention secures true edge-to-edge coherence.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.