Generative Vision Interview Questions, issue 14, Jun 22, 2026
The Shared FFN Trap
Senior AI Engineer interview at Stability AI, and the interviewer asks:
“You’re building an MMDiT. Your teammate wants to share the feed-forward weights across text and image tokens, single-stream, because it’s simpler and cheaper. What are you actually giving up?”
Don’t say: “Nothing much, the attention layer still lets them interact.”
Why saving parameters by merging text and image streams quietly destroys representational specialization, and how separating the "talking" from the "thinking" unlocks elite performance.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.