LLM System Design Interview, issue 12, Nov 10, 2025

The MoE Collapse Trap

Senior ML Engineer interview at Google DeepMind, and the interviewer asks:

โ€œYou have just launched a new Mixture of Experts (MoE) training run. After a few thousand steps, you check the logs and see the validation loss has flatlined. What is the ๐ฆ๐จ๐ฌ๐ญ ๐ฅ๐ข๐ค๐ž๐ฅ๐ฒ ๐œ๐š๐ฎ๐ฌ๐ž specific to an MoE, and how do you fix it?โ€

Donโ€™t say: โ€œMy learning rate is too high,โ€

When your Mixture of Experts stops learning, it's not the optimizer - it's the router. How expert starvation silently turns a 500B model into a 50B one (and how to fix it).

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.