Advanced NLP Interview Questions, issue 22, Dec 27, 2025
The Inter-Annotator Agreement Trap
Senior ML Engineer interview at Google, and the interviewer asks:
“We’re building a toxicity detection dataset where only 1% of comments are actually toxic. We hired two annotators. Their inter-annotator agreement is 99%. Are we good to go?”
Why accuracy collapses under class imbalance - and why Cohen’s Kappa is the metric that actually matters.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.