Computer Vision Interview Questions, issue 15, Jan 16, 2026
The Multimodal Geometry Trap
Senior AI Engineer interview at Meta, and the interviewer asks:
“We are building a Multimodal LLM like LLaVA. We need to feed the frozen CLIP image embeddings into our Language Model. Should we use the final [CLS] token?”
How contrastive pretraining collapses spatial information - and why LLaVA-style models must use penultimate patch embeddings.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.