Computer Vision Interview Questions, issue 14, Jan 15, 2026

The Attention vs MLP Responsibility Trap

AI Researcher interview at OpenAI, and the interviewer asks:

β€œWe know 𝘚𝘦𝘭𝘧-𝘈𝘡𝘡𝘦𝘯𝘡π˜ͺ𝘰𝘯 handles the context between tokens. So, why do we burn ~60% of our parameter budget on the π˜—π˜°π˜΄π˜ͺ𝘡π˜ͺ𝘰𝘯-𝘸π˜ͺ𝘴𝘦 π˜”π˜“π˜— 𝘭𝘒𝘺𝘦𝘳𝘴? What is the MLP actually doing?”

Why attention handles communication, but MLPs do the real computation in modern vision transformers.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.