Generative Vision Interview Questions, issue 13, Jun 21, 2026
The Cross-Attention Trap
Senior AI Engineer interview at Midjourney, and the interviewer asks:
“You switched your text-to-image model from cross-attention to joint attention. Walk me through what actually changes about how text and image tokens talk to each other, and what specific failure that fixes.”
Don’t say: “Joint attention is better because Stable Diffusion 3 uses it.”
Why traditional text conditioning quietly bleeds attributes across your generated image, and the bidirectional sequence fusion that finally fixes compositional binding.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.