LLM Inference Interview Questions, issue 15, Aug 14, 2026

The Abstention Collapse Trap

Staff AI Engineer interview at OpenAI, and the interviewer asks:

You’re running RL to teach a model tool use. Your correctness reward gives partial credit, overlap on tool names, then parameter names, then parameter values. What does the policy learn to exploit before it learns to call tools correctly?

Don’t say: It might reward-hack, so I’d tune the coefficients.

How rewarding parameter overlap silently destroys your agent's ability to say "I don't know", and why you must grade the outcome, not the trace.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.