LLM Agents Interview Questions, issue 13, Mar 7, 2026
The Reward Model Scaling Trap
Senior AI Engineer interview at Anthropic, and the interviewer asks:
“Our RLHF pipeline on an 8B policy model is flatlining on reasoning tasks. A junior engineer wants to scale the Reward Model (RM) from 8B to 70B parameters to get better preference signals. Do you approve the compute budget?”
If your RLHF pipeline stalls on reasoning, scaling the reward model only amplifies generic preference signals instead of fixing the missing task-specific supervision.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.