LLM System Design Interview, issue 71, Sep 15, 2026
The Preference Eval Trap
Senior LLM Post-Training Engineer interview at OpenAI, and the interviewer asks:
“Your new SFT mix lifts head-to-head win rates by 15 points, but MMLU, GSM8K, and your internal capability suite are completely flat. Leadership wants to ship it as a capability gain. What do you tell them?”
Don’t say: “The preference eval is the one users actually feel, so ship it.”
Why crowd raters quietly reward markdown over correctness, and how Bradley-Terry covariate modeling reveals whether your +15% win rate is real or just a system prompt in disguise.
The full answer, with the mechanism and the arithmetic, is free on Substack.