LLM System Design Interview, issue 60, Sep 4, 2026

The Micro-Batch Scaling Trap

Principal ML Infrastructure Engineer interview at NVIDIA, and the interviewer asks:

Your pipeline-parallel run is showing 60% GPU idle time. Your junior says ‘just increase the micro-batches.’ Is he right?

Don’t say: Yes, more micro-batches means fewer pipeline bubbles.

The hidden cliff where amortizing pipeline idle time destroys your arithmetic intensity, and why real training speed comes from stage balancing, not cranking m to infinity.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.