Midjourney interview traps

13 traps set in interviews at Midjourney. Each one: the interviewer's question, the answer most candidates give, and the mechanism that breaks it.

Roles: Senior ML Engineer (7), Senior AI Engineer (5), Generative AI Engineer (1).

From Generative Vision Interview Questions.

Each trap is set in an interview at Midjourney. AI Interview Prep isn't affiliated with Midjourney.

  1. Senior ML Engineer interview at Midjourney

    The Parameter Shrink Trap

    Your text-to-image inference costs are killing the business. Your first instinct is to shrink the student model. Walk me through why that’s the wrong move, and what you’d cut instead.

    Generative Vision #23
    Jul 2, 2026

  2. Generative AI Engineer interview at Midjourney

    The DreamBooth Trap

    You DreamBooth a model on 5 photos of a client’s product. It generates that product perfectly, but now everything else it draws looks broken. Why, and how do you fix it without re-collecting data?

    Generative Vision #22
    Jun 30, 2026

  3. Senior ML Engineer interview at Midjourney

    The Inference Translation Trick

    Your text-to-image model crushes it on internal evals, but users complain the outputs look flat. They type ‘a teddy bear reading a book’ and get garbage. The weights are fine. What’s actually broken?

    Generative Vision #21
    Jun 29, 2026

  4. Senior ML Engineer interview at Midjourney

    The SFT Misdiagnosis Trap

    Your text-to-image model renders the prompt correctly, right objects, right layout, but every output looks flat and amateur. Walk me through your fix.

    Generative Vision #19
    Jun 27, 2026

  5. Senior ML Engineer interview at Midjourney

    The Perceived Noise Paradox

    You trained a diffusion model that’s gorgeous at 256px. You scale it to 1024px, keep the exact same noise schedule, and the outputs get quietly worse. Same architecture, same loss, same schedule. What broke?

    Generative Vision #17
    Jun 25, 2026

  6. Senior ML Engineer interview at Midjourney

    The Mid-Noise Paradox

    Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?

    Generative Vision #16
    Jun 24, 2026

  7. Senior ML Engineer interview at Midjourney

    The Resolution Extrapolation Trap

    Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?

    Generative Vision #15
    Jun 23, 2026

  8. Senior AI Engineer interview at Midjourney

    The Cross-Attention Trap

    You switched your text-to-image model from cross-attention to joint attention. Walk me through what actually changes about how text and image tokens talk to each other, and what specific failure that fixes.

    Generative Vision #13
    Jun 21, 2026

  9. Senior ML Engineer interview at Midjourney

    The Receptive Field Illusion

    Your teammate wants to ship a U-Net for your new high-res image model because ‘convolutions capture both local and global features.’ You disagree. Defend it.

    Generative Vision #10
    Jun 18, 2026

  10. Senior AI Engineer interview at Midjourney

    The Spatial Addressing Paradox

    Your text-to-image model nails single subjects, but prompt it with ‘a brown teddy bear next to a white wall’ and the colors bleed across the entire frame. An intern on your team says ‘add more training data.’ Why is that wrong, and where is this actually breaking?

    Generative Vision #9
    Jun 17, 2026

  11. Senior AI Engineer interview at Midjourney

    The Mode Ascent Trap

    Your score estimator loss is perfectly converged after 400 hours on an A100 cluster. But during inference, deterministically following the gradient generates the exact same 3 hyper-average images on repeat. Why?

    Generative Vision #5
    Jun 12, 2026

  12. Senior AI Engineer interview at Midjourney

    The SNR Collapse Trap

    Images are discrete RGB values from 0 to 255. Diffusion math assumes we sample from a standard normal distribution. If a data engineer feeds raw 0-255 pixel tensors directly into the training pipeline without continuous float scaling, how does this mathematically break the variance-preserving nature of the forward process?

    Generative Vision #4
    Jun 11, 2026

  13. Senior AI Engineer interview at Midjourney

    The Noise Schedule Trap

    Your diffusion model generates photorealistic textures, but the global shapes are completely mangled ( for example three-headed teddy bears ). The architecture is flawless. What phase of your forward noise schedule (𝛽_𝑡) is failing, and why?

    Generative Vision #1
    Jun 8, 2026

Get the next one

Free on Substack. Unsubscribe in one click.