LLM System Design Interview, issue 50, May 13, 2026

The Rejection Sampling Paradox

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

You deployed a 70B target model with a 1B draft model for speculative decoding. Accuracy is identical, but your expected 2x speedup is sitting at exactly 0%. Why?

Why your expected 2x inference speedup is sitting at exactly 0%, and how domain-specific alignment and dynamic lookahead actually fix the speculative decoding bottleneck.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.