Computer Vision Interview Questions, issue 18, Jan 19, 2026

The Compositionality Trap

Senior Computer Vision Engineer interview at Google DeepMind, and the interviewer asks:

β€œOur YOLO model has 99% mAP ( Mean Average Precision ) on π˜—π˜¦π˜°π˜±π˜­π˜¦ and 𝘍π˜ͺ𝘳𝘦 𝘏𝘺π˜₯𝘳𝘒𝘯𝘡𝘴 individually. But in production, we saw a person sitting on a fire hydrant, and the model didn’t flag it as anomalous. It just saw two boxes. Why did we fail, and how do you fix it?”

Why 99% Mean Average Precision fails in the real world, and why bounding boxes can’t reason about relationships.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.