Computer Vision Interview Questions, issue 11, Jan 12, 2026
The CLIP Prompt Variance Trap
Senior Computer Vision Engineer interview at OpenAI, and the interviewer asks:
โWe just deployed a CLIP model for zero-shot classification. Weโre feeding in raw class names like ๐ฅ๐ฐ๐จ or ๐ฑ๐ญ๐ข๐ฏ๐ฆ as text prompts. The accuracy is shaky and the variance is high. Without retraining a single parameter, ๐ก๐จ๐ฐ ๐๐จ ๐ฒ๐จ๐ฎ ๐๐ข๐ฑ ๐ญ๐ก๐ ๐ฌ๐ญ๐๐๐ข๐ฅ๐ข๐ญ๐ฒ ๐๐ง๐ ๐๐จ๐จ๐ฌ๐ญ ๐๐ฆ๐๐ ๐๐๐๐ญ ๐๐๐๐ฎ๐ซ๐๐๐ฒ?โ
Why single-text prompts are noisy estimates in high-dimensional spaceโand how centroid stabilization fixes zero-shot accuracy.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.