LLM System Design Interview, issue 25, Nov 23, 2025
The Leaderboard Illusion
AI Engineer interview at Anthropic, and the interviewer asks:
“Our model is lagging on the Chatbot Arena leaderboard. We need to boost our ELO score by 50 points next release to match GPT-5. How do you adjust the post-training pipeline to make this happen?”
Don’t say: “We need to improve our reasoning capabilities on hard math benchmarks like MATH-500 or GPQA.”
Why chasing Chatbot Arena ELO degrades your model, burns compute, and misaligns you from real customers, and what a senior AI engineer should optimize instead.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.