Machine Learning System Design Interview, issue 50, Jun 7, 2026
The Delayed Reward Illusion
Senior ML Engineer interview at Netflix, and the interviewer asks:
“We are launching a new recommendation variant. Under what infrastructure constraints and business risks is a Multi-Armed Bandit (MAB) the wrong choice over a basic A/B test?”
Why forcing real-time state management on weeks-long conversion cycles causes massive system lag, and the infrastructure constraints that make Bandits the wrong choice.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.