Computer Vision Interview Questions, issue 23, Jan 24, 2026

The Flamingo Architecture Trap

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

We have a 70B parameter LLM. We need it to ‘see’ images. But here’s the constraint: We have zero budget to fine-tune the 70B weights, and we can’t afford to destroy the model’s existing reasoning capabilities.

The exact components you need to add vision to a frozen LLM - without paying the fine-tuning cost.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.