Generative Vision Interview Questions, issue 15, Jun 23, 2026

The Resolution Extrapolation Trap

Senior ML Engineer interview at Midjourney, and the interviewer asks:

Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?

Don’t say: The model didn’t see high-res data, so I’ll fine-tune at 1024.

Why bumping up image resolution silently destroys your DiT's spatial awareness, and how to rescale your coordinate grid to eliminate repeated limbs without touching a single weight.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.