Advanced NLP Interview Questions, issue 20, Dec 26, 2025
The Quantization Gradient Trap
Senior AI Engineer interview at Meta, and the interviewer asks:
โWe need to switch to ๐๐ฎ๐๐ง๐ญ๐ข๐ณ๐๐ญ๐ข๐จ๐ง ๐๐ฐ๐๐ซ๐ ๐๐ซ๐๐ข๐ง๐ข๐ง๐ (๐๐๐) because post-training quantization is tanking our accuracy. But the rounding operation (Float -\> Int8) is a step function with a derivative of zero. How do you actually backpropagate gradients through it to update the weights?โ
Why rounding kills backpropagation - and how STE โliesโ to the optimizer to keep QAT training alive.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.