LLM Inference Interview Questions, issue 6, Aug 4, 2026

The Multimodal Perception Trap

Senior AI Engineer interview at Google, and the interviewer asks:

Your web agent uses set-of-marks, you screenshot the page, draw numbered boxes on every element, and let the VLM click by number. It works in your demo but in production the model keeps clicking box 41 when it meant box 14, and it ignores half the page. The VLM is state-of-the-art and multimodal. So why is it failing?

Don’t say: We need a bigger/better multimodal model.

Why dropping frontier VLMs onto dense web UIs is a hidden trap, and how foundation labs use cheap synthetic data to turn a prompting party trick into reliable execution.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.