Computer Vision Interview Questions, issue 25, Jan 26, 2026
The Contrastive Shortcut Trap
Computer Vision Engineer interview at OpenAI, and the interviewer asks:
βWe are building a π‘π¦π³π°-ππ©π°π΅ πππ’π΄π΄πͺπ§πͺπ¦π³. We have the budget for a standard CLIP architecture. Why should we burn 25% more VRAM adding a ππ¦π―π¦π³π’π΅πͺπ·π¦ ππ¦π€π°π₯π¦π³ (ππ°ππ’) if we donβt need to generate captions?β
Why CLIP learns just enough to pass, and how a generative decoder forces the encoder to stop being lazy.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.