LLM Agents Interview Questions, issue 17, Mar 12, 2026

The Weak-Negative Data Trap

Senior ML Engineer interview at Anthropic, and the interviewer asks:

We need to train a rigorous LLM-as-a-Judge, but we have absolutely zero human-labeled preference pairs. How do you procedurally generate ‘rejected’ responses that are nuanced enough to actually train a robust evaluator?

Don’t say: Just prompt a smaller, weaker model to write a bad answer, or randomly inject formatting errors into a good response.

Training an evaluator on sloppy responses creates a grammar detector, not a reasoning judge capable of catching articulate hallucinations.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.