Computer Vision Interview Questions

Convolutions, detection and the geometry underneath.

25 traps, Dec 2025 to Jan 2026. Complete.

Set in interviews at OpenAI (11), Google DeepMind (7), Google (2) and Meta (2).

The retriever for Computer Vision Interview Questions, drawn at a desk

Each trap: the interviewerโ€™s question, the answer most candidates give, and the mechanism that breaks it. The full answers are on Substack.

  1. Computer Vision Engineer interview at OpenAI

    The Contrastive Shortcut Trap

    โ€œWe are building a ๐˜ก๐˜ฆ๐˜ณ๐˜ฐ-๐˜š๐˜ฉ๐˜ฐ๐˜ต ๐˜Š๐˜ญ๐˜ข๐˜ด๐˜ด๐˜ช๐˜ง๐˜ช๐˜ฆ๐˜ณ. We have the budget for a standard CLIP architecture. Why should we burn 25% more VRAM adding a ๐˜Ž๐˜ฆ๐˜ฏ๐˜ฆ๐˜ณ๐˜ข๐˜ต๐˜ช๐˜ท๐˜ฆ ๐˜‹๐˜ฆ๐˜ค๐˜ฐ๐˜ฅ๐˜ฆ๐˜ณ (๐˜Š๐˜ฐ๐˜Š๐˜ข) if we donโ€™t need to generate captions?โ€

    Computer Vision #25
    Jan 26, 2026

  2. Senior AI Engineer interview at Google DeepMind

    The Signal-to-Noise Trap

    โ€œOur competitor just trained a VLM on 6 billion image-text pairs. We only have the compute budget for 700k images. How do we beat them?โ€

    Computer Vision #24
    Jan 25, 2026

  3. Senior AI Engineer interview at Google DeepMind

    The Flamingo Architecture Trap

    โ€œWe have a 70B parameter LLM. We need it to โ€˜seeโ€™ images. But hereโ€™s the constraint: We have zero budget to fine-tune the 70B weights, and we canโ€™t afford to destroy the modelโ€™s existing reasoning capabilities.โ€

    Computer Vision #23
    Jan 24, 2026

  4. final-round Computer Vision Engineer interview at OpenAI

    The Interactive Segmentation Trap

    โ€œYour user clicks here. What mask does your model output?โ€

    Computer Vision #22
    Jan 23, 2026

  5. Senior Robotics Engineer interview at NVIDIA

    The Data Scaling Trap

    โ€œWe need a robot to open any drawer in any userโ€™s home. We cannot pre-train it on every possible handle shape. How do you build this?โ€

    Computer Vision #21
    Jan 22, 2026

  6. Senior Computer Vision Engineer interview at OpenAI

    The Low-Contrast Bias Trap

    โ€œOur production FaceID model has a 12% higher error rate on darker skin tones. We audited the training data and it is perfectly balanced (50/50 split). We retrained from scratch. The error persists. Why?โ€

    Computer Vision #20
    Jan 21, 2026

  7. Senior Computer Vision Engineer interview at Google DeepMind

    The Fine-Grained Invariance Trap

    โ€œWe need to classify 10,000 distinct car models (Make, Model, Year) for a demographics study. How do you build the model?โ€

    Computer Vision #19
    Jan 20, 2026

  8. Senior Computer Vision Engineer interview at Google DeepMind

    The Compositionality Trap

    โ€œOur YOLO model has 99% mAP ( Mean Average Precision ) on ๐˜—๐˜ฆ๐˜ฐ๐˜ฑ๐˜ญ๐˜ฆ and ๐˜๐˜ช๐˜ณ๐˜ฆ ๐˜๐˜บ๐˜ฅ๐˜ณ๐˜ข๐˜ฏ๐˜ต๐˜ด individually. But in production, we saw a person sitting on a fire hydrant, and the model didnโ€™t flag it as anomalous. It just saw two boxes. Why did we fail, and how do you fix it?โ€

    Computer Vision #18
    Jan 19, 2026

  9. Senior AI Engineer interview at OpenAI

    The Counting Hallucination Trap

    โ€œOur VLM constantly hallucinates object counts in crowded images. It says โ€˜8 peopleโ€™ when there are only 5. We have zero budget for new data collection. How do you fix this?โ€

    Computer Vision #17
    Jan 18, 2026

  10. Senior AI Engineer interview at OpenAI

    The Contrastive Hard Negative Trap

    โ€œOur CLIP model keeps confusing Golden Retrievers with Yellow Labs. To fix it, weโ€™re going to manually curate hard negative batches, forcing these similar breeds into the same training step. Good idea?โ€

    Computer Vision #16
    Jan 17, 2026

  11. Senior AI Engineer interview at Meta

    The Multimodal Geometry Trap

    โ€œWe are building a Multimodal LLM like LLaVA. We need to feed the frozen CLIP image embeddings into our Language Model. Should we use the final [CLS] token?โ€

    Computer Vision #15
    Jan 16, 2026

  12. AI Researcher interview at OpenAI

    The Attention vs MLP Responsibility Trap

    โ€œWe know ๐˜š๐˜ฆ๐˜ญ๐˜ง-๐˜ˆ๐˜ต๐˜ต๐˜ฆ๐˜ฏ๐˜ต๐˜ช๐˜ฐ๐˜ฏ handles the context between tokens. So, why do we burn ~60% of our parameter budget on the ๐˜—๐˜ฐ๐˜ด๐˜ช๐˜ต๐˜ช๐˜ฐ๐˜ฏ-๐˜ธ๐˜ช๐˜ด๐˜ฆ ๐˜”๐˜“๐˜— ๐˜ญ๐˜ข๐˜บ๐˜ฆ๐˜ณ๐˜ด? What is the MLP actually doing?โ€

    Computer Vision #14
    Jan 15, 2026

  13. Senior Computer Vision Engineer interview at Google DeepMind

    The Generalization Gap Trap

    โ€œWe use heavy data augmentation (Color Jitter, 30ยฐ Rotations) during training to improve robustness. Why do we strictly disable these during validation? Doesnโ€™t that break the rule that ๐˜›๐˜ณ๐˜ข๐˜ช๐˜ฏ ๐˜ข๐˜ฏ๐˜ฅ ๐˜›๐˜ฆ๐˜ด๐˜ต ๐˜ฅ๐˜ช๐˜ด๐˜ต๐˜ณ๐˜ช๐˜ฃ๐˜ถ๐˜ต๐˜ช๐˜ฐ๐˜ฏ๐˜ด ๐˜ด๐˜ฉ๐˜ฐ๐˜ถ๐˜ญ๐˜ฅ ๐˜ฎ๐˜ข๐˜ต๐˜ค๐˜ฉ?โ€

    Computer Vision #13
    Jan 14, 2026

  14. Senior AI Engineer interview at Google DeepMind

    The Large Batch Generalization Trap

    โ€œWe just scaled our infrastructure to 4x our batch size (256 to 1024) to speed up training. We followed the ๐˜“๐˜ช๐˜ฏ๐˜ฆ๐˜ข๐˜ณ ๐˜š๐˜ค๐˜ข๐˜ญ๐˜ช๐˜ฏ๐˜จ ๐˜™๐˜ถ๐˜ญ๐˜ฆ and multiplied our ๐˜“๐˜ฆ๐˜ข๐˜ณ๐˜ฏ๐˜ช๐˜ฏ๐˜จ ๐˜™๐˜ข๐˜ต๐˜ฆ by 4. But our test accuracy still degraded. What fundamental property of SGD did we accidentally kill?โ€

    Computer Vision #12
    Jan 13, 2026

  15. Senior Computer Vision Engineer interview at OpenAI

    The CLIP Prompt Variance Trap

    โ€œWe just deployed a CLIP model for zero-shot classification. Weโ€™re feeding in raw class names like ๐˜ฅ๐˜ฐ๐˜จ or ๐˜ฑ๐˜ญ๐˜ข๐˜ฏ๐˜ฆ as text prompts. The accuracy is shaky and the variance is high. Without retraining a single parameter, ๐ก๐จ๐ฐ ๐๐จ ๐ฒ๐จ๐ฎ ๐Ÿ๐ข๐ฑ ๐ญ๐ก๐ž ๐ฌ๐ญ๐š๐›๐ข๐ฅ๐ข๐ญ๐ฒ ๐š๐ง๐ ๐›๐จ๐จ๐ฌ๐ญ ๐ˆ๐ฆ๐š๐ ๐ž๐๐ž๐ญ ๐š๐œ๐œ๐ฎ๐ซ๐š๐œ๐ฒ?โ€

    Computer Vision #11
    Jan 12, 2026

  16. Computer Vision Engineer interview at Meta

    The Early vs Slow Fusion Trap

    โ€œWeโ€™re debating between ๐˜Œ๐˜ข๐˜ณ๐˜ญ๐˜บ ๐˜๐˜ถ๐˜ด๐˜ช๐˜ฐ๐˜ฏ and ๐˜š๐˜ญ๐˜ฐ๐˜ธ ๐˜๐˜ถ๐˜ด๐˜ช๐˜ฐ๐˜ฏ for our new video understanding model. Everyone knows ๐˜š๐˜ญ๐˜ฐ๐˜ธ ๐˜๐˜ถ๐˜ด๐˜ช๐˜ฐ๐˜ฏ captures motion better, but what is the specific computational consequence of maintaining that temporal dimension through multiple layers that kills our training budget?โ€

    Computer Vision #10
    Jan 11, 2026

  17. Senior Computer Vision Engineer interview at Amazon Fulfillment Technologies & Robotics

    The Tiny Object Trap

    โ€œWe need to detect tiny, 3mm micro-fractures on a fast-moving assembly line. You suggested ๐…๐š๐ฌ๐ญ๐ž๐ซ ๐‘-๐‚๐๐ over ๐˜๐Ž๐‹๐Ž. Why does the ๐‘๐ž๐ ๐ข๐จ๐ง ๐๐ซ๐จ๐ฉ๐จ๐ฌ๐š๐ฅ ๐๐ž๐ญ๐ฐ๐จ๐ซ๐ค (๐‘๐๐) specifically help with small objects, even though it kills our inference speed?โ€

    Computer Vision #9
    Jan 10, 2026

  18. Senior Computer Vision Engineer interview at OpenAI

    The Zero-Padding Distribution Trap

    โ€œWe use Zero-Padding to maintain feature map dimensions (e.g., 32x32). But from a signal processing perspective, why is injecting zeros at the borders dangerous for your modelโ€™s statistical distribution?โ€

    Computer Vision #8
    Jan 9, 2026

  19. Senior Computer Vision Engineer interview at OpenAI

    The Receptive Field Trap

    โ€œIn VGGNet, we replace a single 7x7 convolution with a stack of three 3x3 convolutions. Why?โ€

    Computer Vision #7
    Jan 8, 2026

  20. Senior MLE Engineer interview at Google DeepMind

    The Model Capacity Trap

    โ€œOur new foundation model is overfitting severely on the training set. Should we cut the hidden dimension size from 4096 to 1024 to limit its capacity?โ€

    Computer Vision #6
    Jan 7, 2026

  21. Senior ML Engineer interview at OpenAI

    The Dead ReLU Trap

    โ€œHow do you fix this?โ€

    Computer Vision #5
    Jan 6, 2026

  22. Senior Computer Vision Engineer interview at Google

    The L1 vs L2 Geometry Trap

    โ€œWeโ€™re building a similarity search for a new dataset. If I arbitrarily rotate the feature space by 45 degrees, which distance metric falls apart: ๐‹1 ๐จ๐ซ ๐‹2? And what does that tell you about our feature engineering strategy?โ€

    Computer Vision #4
    Jan 5, 2026

  23. Machine Learning Engineer interview at OpenAI

    The Low Initial Loss Trap

    โ€œYou kick off training for a Softmax classifier on CIFAR-10 (10 classes). In the very first iteration, your loss reads 0.05. Is this good news?โ€

    Computer Vision #3
    Jan 4, 2026

  24. Senior Computer Vision Engineer interview at Google

    The Redundant Data Trap

    โ€œWe trained a high-capacity ResNet on 500k images, but itโ€™s still overfitting. My Product Manager wants to spend $20k to label another 500k random images scraped from the same source. Do you approve the budget?โ€

    Computer Vision #2
    Jan 3, 2026

  25. Senior Computer Vision Engineer interview at Tesla

    The Translation Equivariance Efficiency Trap

    โ€œWe all know ๐‚๐๐๐ฌ are translation equivariant. But why exactly does that property make them exponentially more data-efficient than a ๐…๐ฎ๐ฅ๐ฅ๐ฒ ๐‚๐จ๐ง๐ง๐ž๐œ๐ญ๐ž๐ ๐ง๐ž๐ญ๐ฐ๐จ๐ซ๐ค for processing high-res images?โ€

    Computer Vision #1
    Dec 31, 2025

Get the next one

Free on Substack. Unsubscribe in one click.