Generative Vision Interview Questions, issue 16, Jun 24, 2026

The Mid-Noise Paradox

Senior ML Engineer interview at Midjourney, and the interviewer asks:

Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?

Don’t say: The model didn’t see high-res data, so I’ll fine-tune at 1024.

Why 2x-ing your training compute won't fix blurry images, and how switching to a logit-normal curriculum forces your model to fight the structural war where it actually matters.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.