LLM System Design Interview, issue 60, Sep 4, 2026
The Micro-Batch Scaling Trap
Principal ML Infrastructure Engineer interview at NVIDIA, and the interviewer asks:
“Your pipeline-parallel run is showing 60% GPU idle time. Your junior says ‘just increase the micro-batches.’ Is he right?”
Don’t say: “Yes, more micro-batches means fewer pipeline bubbles.”
The hidden cliff where amortizing pipeline idle time destroys your arithmetic intensity, and why real training speed comes from stage balancing, not cranking m to infinity.
The full answer, with the mechanism and the arithmetic, is free on Substack.