Computer Vision Interview Questions, issue 17, Jan 18, 2026

The Counting Hallucination Trap

Senior AI Engineer interview at OpenAI, and the interviewer asks:

Our VLM constantly hallucinates object counts in crowded images. It says ‘8 people’ when there are only 5. We have zero budget for new data collection. How do you fix this?

Why caption-only supervision lets VLMs hallucinate counts, and how forcing spatial proof fixes it.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.