Generative Vision Interview Questions, issue 23, Jul 2, 2026
The Parameter Shrink Trap
Senior ML Engineer interview at Midjourney, and the interviewer asks:
“Your text-to-image inference costs are killing the business. Your first instinct is to shrink the student model. Walk me through why that’s the wrong move, and what you’d cut instead.”
Don’t say: “Smaller model, fewer parameters, faster inference, just distill it down like DistilBERT.”
Why cutting network width won't save your text-to-image inference costs, and how the progressive distillation trick collapses latency without sacrificing the canvas.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.