Advanced NLP Interview Questions, issue 20, Dec 26, 2025

The Quantization Gradient Trap

Senior AI Engineer interview at Meta, and the interviewer asks:

โ€œWe need to switch to ๐๐ฎ๐š๐ง๐ญ๐ข๐ณ๐š๐ญ๐ข๐จ๐ง ๐€๐ฐ๐š๐ซ๐ž ๐“๐ซ๐š๐ข๐ง๐ข๐ง๐  (๐๐€๐“) because post-training quantization is tanking our accuracy. But the rounding operation (Float -\> Int8) is a step function with a derivative of zero. How do you actually backpropagate gradients through it to update the weights?โ€

Why rounding kills backpropagation - and how STE โ€œliesโ€ to the optimizer to keep QAT training alive.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.