Generative Vision Interview Questions, issue 16, Jun 24, 2026
The Mid-Noise Paradox
Senior ML Engineer interview at Midjourney, and the interviewer asks:
“Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?”
Don’t say: “The model didn’t see high-res data, so I’ll fine-tune at 1024.”
Why 2x-ing your training compute won't fix blurry images, and how switching to a logit-normal curriculum forces your model to fight the structural war where it actually matters.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.