Advanced Reinforcement Learning Interview Questions, issue 22, Feb 17, 2026
The Information Density Trap
Senior RLHF Engineer interview at OpenAI, and the interviewer asks:
“We have a \$50k budget for human labeling. We need a reward model for ‘helpfulness.’ Do we pay humans to score responses on a 1-10 scale, or rank pairs (A \> B)?”
Maximizing numeric “richness” with 1–10 scores backfires because inconsistent human baselines corrupt the signal before the model ever sees it.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.