LLM System Design Interview, issue 71, Sep 15, 2026

The Preference Eval Trap

Senior LLM Post-Training Engineer interview at OpenAI, and the interviewer asks:

Your new SFT mix lifts head-to-head win rates by 15 points, but MMLU, GSM8K, and your internal capability suite are completely flat. Leadership wants to ship it as a capability gain. What do you tell them?

Don’t say: The preference eval is the one users actually feel, so ship it.

Why crowd raters quietly reward markdown over correctness, and how Bradley-Terry covariate modeling reveals whether your +15% win rate is real or just a system prompt in disguise.

The full answer, with the mechanism and the arithmetic, is free on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.