Computer Vision Interview Questions, issue 11, Jan 12, 2026

The CLIP Prompt Variance Trap

Senior Computer Vision Engineer interview at OpenAI, and the interviewer asks:

โ€œWe just deployed a CLIP model for zero-shot classification. Weโ€™re feeding in raw class names like ๐˜ฅ๐˜ฐ๐˜จ or ๐˜ฑ๐˜ญ๐˜ข๐˜ฏ๐˜ฆ as text prompts. The accuracy is shaky and the variance is high. Without retraining a single parameter, ๐ก๐จ๐ฐ ๐๐จ ๐ฒ๐จ๐ฎ ๐Ÿ๐ข๐ฑ ๐ญ๐ก๐ž ๐ฌ๐ญ๐š๐›๐ข๐ฅ๐ข๐ญ๐ฒ ๐š๐ง๐ ๐›๐จ๐จ๐ฌ๐ญ ๐ˆ๐ฆ๐š๐ ๐ž๐๐ž๐ญ ๐š๐œ๐œ๐ฎ๐ซ๐š๐œ๐ฒ?โ€

Why single-text prompts are noisy estimates in high-dimensional spaceโ€”and how centroid stabilization fixes zero-shot accuracy.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.