LLM System Design Interview, issue 40, May 3, 2026

The Expert Capacity Paradox

Senior ML Engineer interview at Google DeepMind, and the interviewer asks:

Sending the exact same prompt yields slightly different outputs depending on the time of day.

Don’t say: It must be a KV cache corruption issue. Since Temperature 0 is mathematically deterministic, the continuous batching engine must be leaking state, or there’s an alignment bug in your padding masks.

Why batched MoE inference silently hallucinates bugs even at zero temperature, and how to trade FLOPs for drop-free routing to restore mathematical determinism.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.