Generative Vision Interview Questions, issue 19, Jun 27, 2026

The SFT Misdiagnosis Trap

Senior ML Engineer interview at Midjourney, and the interviewer asks:

Your text-to-image model renders the prompt correctly, right objects, right layout, but every output looks flat and amateur. Walk me through your fix.

Don’t say: I’d collect more training data and keep training until quality improves.

The hidden reason teams quietly waste pre-training scale compute on simple aesthetic fixes, and how to definitively map visual flaws to the exact post-training lever they require.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.