Computer Vision Interview Questions, issue 14, Jan 15, 2026
The Attention vs MLP Responsibility Trap
AI Researcher interview at OpenAI, and the interviewer asks:
βWe know ππ¦ππ§-ππ΅π΅π¦π―π΅πͺπ°π― handles the context between tokens. So, why do we burn ~60% of our parameter budget on the ππ°π΄πͺπ΅πͺπ°π―-πΈπͺπ΄π¦ πππ ππ’πΊπ¦π³π΄? What is the MLP actually doing?β
Why attention handles communication, but MLPs do the real computation in modern vision transformers.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.