Advanced NLP Interview Questions, issue 5, Dec 12, 2025

The Speculative Decoding Illusion

Machine Learning Engineer interview at OpenAI, and the interviewer asks:

We need to optimize inference for batch size 128. Should we use Speculative Decoding?

You think batch=128 makes speculative decoding useless. The reality? Your model is secretly dying of memory starvation.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.