Generative Vision Interview Questions, issue 10, Jun 18, 2026

The Receptive Field Illusion

Senior ML Engineer interview at Midjourney, and the interviewer asks:

Your teammate wants to ship a U-Net for your new high-res image model because ‘convolutions capture both local and global features.’ You disagree. Defend it.

Don’t say: Transformers are just better, everyone uses DiT now.

Why stacking convolutions to build global context quietly turns distant features into mush, and how trading inductive bias for self-attention secures true edge-to-edge coherence.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.