LLM System Design Interview, issue 25, Nov 23, 2025

The Leaderboard Illusion

AI Engineer interview at Anthropic, and the interviewer asks:

Our model is lagging on the Chatbot Arena leaderboard. We need to boost our ELO score by 50 points next release to match GPT-5. How do you adjust the post-training pipeline to make this happen?

Don’t say: We need to improve our reasoning capabilities on hard math benchmarks like MATH-500 or GPQA.

Why chasing Chatbot Arena ELO degrades your model, burns compute, and misaligns you from real customers, and what a senior AI engineer should optimize instead.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.