Machine Learning System Design Interview, issue 3, Nov 25, 2025

The Gradient Drowning Trap

Senior Machine Learning Engineer interview at Google DeepMind, and the interviewer asks:

β€œWe have a 1:1000 class imbalance for fraud detection. We applied 𝘀𝘭𝘒𝘴𝘴_𝘸𝘦π˜ͺ𝘨𝘩𝘡𝘴 to the 𝐂𝐫𝐨𝐬𝐬-𝐄𝐧𝐭𝐫𝐨𝐩𝐲 loss, but the model is still missing the hard edge cases. What do we do?”

The hidden gradient dynamics that SMOTE, class weights, and oversampling can’t fix.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.