Advanced NLP Interview Questions, issue 5, Dec 12, 2025
The Speculative Decoding Illusion
Machine Learning Engineer interview at OpenAI, and the interviewer asks:
“We need to optimize inference for batch size 128. Should we use Speculative Decoding?”
You think batch=128 makes speculative decoding useless. The reality? Your model is secretly dying of memory starvation.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.