LLM System Design Interview, issue 11, Nov 9, 2025
The Alignment Tax
AI Engineer interview at OpenAI, and the interviewer asks:
“You’ve successfully fine-tuned a model with RL. It’s now excellent at following instructions, but it’s become ‘dumber’ at general knowledge and creative writing. What is this phenomenon called, and what specific term would you add to your loss function to prevent this?”
Don’t say: “That’s catastrophic forgetting. We can fix it with a lower learning rate or by mixing in more general-purpose SFT data.”
When RL makes your model obedient but dumb - and how a KL divergence leash keeps it smart, creative, and aligned.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.