Computer Vision Interview Questions, issue 18, Jan 19, 2026
The Compositionality Trap
Senior Computer Vision Engineer interview at Google DeepMind, and the interviewer asks:
βOur YOLO model has 99% mAP ( Mean Average Precision ) on ππ¦π°π±ππ¦ and ππͺπ³π¦ ππΊπ₯π³π’π―π΅π΄ individually. But in production, we saw a person sitting on a fire hydrant, and the model didnβt flag it as anomalous. It just saw two boxes. Why did we fail, and how do you fix it?β
Why 99% Mean Average Precision fails in the real world, and why bounding boxes canβt reason about relationships.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.