Generative Vision Interview Questions, issue 15, Jun 23, 2026
The Resolution Extrapolation Trap
Senior ML Engineer interview at Midjourney, and the interviewer asks:
“Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?”
Don’t say: “The model didn’t see high-res data, so I’ll fine-tune at 1024.”
Why bumping up image resolution silently destroys your DiT's spatial awareness, and how to rescale your coordinate grid to eliminate repeated limbs without touching a single weight.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.