Computer Vision Interview Questions, issue 15, Jan 16, 2026

The Multimodal Geometry Trap

Senior AI Engineer interview at Meta, and the interviewer asks:

We are building a Multimodal LLM like LLaVA. We need to feed the frozen CLIP image embeddings into our Language Model. Should we use the final [CLS] token?

How contrastive pretraining collapses spatial information - and why LLaVA-style models must use penultimate patch embeddings.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.