Computer Vision Interview Questions, issue 25, Jan 26, 2026

The Contrastive Shortcut Trap

Computer Vision Engineer interview at OpenAI, and the interviewer asks:

β€œWe are building a 𝘑𝘦𝘳𝘰-𝘚𝘩𝘰𝘡 𝘊𝘭𝘒𝘴𝘴π˜ͺ𝘧π˜ͺ𝘦𝘳. We have the budget for a standard CLIP architecture. Why should we burn 25% more VRAM adding a 𝘎𝘦𝘯𝘦𝘳𝘒𝘡π˜ͺ𝘷𝘦 π˜‹π˜¦π˜€π˜°π˜₯𝘦𝘳 (𝘊𝘰𝘊𝘒) if we don’t need to generate captions?”

Why CLIP learns just enough to pass, and how a generative decoder forces the encoder to stop being lazy.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.