Generative Vision Interview Questions, issue 19, Jun 27, 2026
The SFT Misdiagnosis Trap
Senior ML Engineer interview at Midjourney, and the interviewer asks:
“Your text-to-image model renders the prompt correctly, right objects, right layout, but every output looks flat and amateur. Walk me through your fix.”
Don’t say: “I’d collect more training data and keep training until quality improves.”
The hidden reason teams quietly waste pre-training scale compute on simple aesthetic fixes, and how to definitively map visual flaws to the exact post-training lever they require.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.