LLM System Design Interview, issue 22, Nov 18, 2025

The Asynchronous Execution Trap

technical Engineer interview at NVIDIA, and the interviewer asks:

An intern excitedly claims they achieved a 1000x speedup on a new matrix multiplication kernel. You look at their script and see they simply wrapped the function call with standard Python timers: 𝘴𝘵𝘢𝘳𝘵 = 𝘵𝘪𝘮𝘦.𝘵𝘪𝘮𝘦() ... 𝘦𝘯𝘥 = 𝘵𝘪𝘮𝘦.𝘵𝘪𝘮𝘦() Why are their results a complete lie?

Why naive Python timers give you “1000× speedups,” why the GPU never actually ran your kernel, and the one missing line (synchronization) that every real ML engineer knows by heart.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.