LLM Agents Interview Questions, issue 13, Mar 7, 2026

The Reward Model Scaling Trap

Senior AI Engineer interview at Anthropic, and the interviewer asks:

Our RLHF pipeline on an 8B policy model is flatlining on reasoning tasks. A junior engineer wants to scale the Reward Model (RM) from 8B to 70B parameters to get better preference signals. Do you approve the compute budget?

If your RLHF pipeline stalls on reasoning, scaling the reward model only amplifies generic preference signals instead of fixing the missing task-specific supervision.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.