LLM System Design Interview, issue 21, Nov 17, 2025

The GRPO Length Trap

Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

We’ve implemented the original DeepSeek GRPO paper to train our new math chatbot. On uncertain queries, the Chain-of-Thought (CoT) is suddenly exploding to 10000 tokens. An engineer on the team says this is great, the model is just thinking harder and learning to backtrack. What’s your diagnosis?

Don’t say: This is a fantastic result! It’s an emergent property. The RL algorithm is clearly working, forcing the model to think more to find the right answer. We can just clip the output at inference time.

How a "good idea" in the DeepSeek objective secretly incentivizes 10000-token failures - and bankrupts your inference budget.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.