LLM System Design Interview, issue 50, May 13, 2026
The Rejection Sampling Paradox
Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
“You deployed a 70B target model with a 1B draft model for speculative decoding. Accuracy is identical, but your expected 2x speedup is sitting at exactly 0%. Why?”
Why your expected 2x inference speedup is sitting at exactly 0%, and how domain-specific alignment and dynamic lookahead actually fix the speculative decoding bottleneck.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.