LLM System Design Interview, issue 73, Sep 17, 2026

The Reasoning Length Trap

Senior ML Engineer interview at Google DeepMind, and the interviewer asks:

You switched to GRPO. Average chain-of-thought length grows every training step. Leadership calls it ‘the model learning to think harder.’ What’s the less flattering explanation, and how do you verify it before your inference bill doubles?

Don’t say: Longer CoT means the model is reasoning more deeply. That’s expected RL behavior, and it’s what DeepSeek R1 showed.

When growing Chain-of-Thought isn't emergent intelligence, but a hidden RL artifact silently diluting negative penalties, and the simple normalization fix that saves your serving budget.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.