LLM System Design Interview, issue 17, Nov 14, 2025
The "Divine Benevolence" Fallacy
Senior ML Engineer interview at OpenAI, and the interviewer asks:
βOne of your junior researchers is burning compute time trying to build a theoretical proof for why ππ°π’πππ outperforms standard ππππ in your new model. Your pre-training deadline is in 48 hours. How do you handle this?β
Donβt say: βIβd encourage their curiosity. Iβd ask them to time-box the research to one more day and present their findings. Understanding the why is key to long-term innovation.β
SwiGLU works. The ablations say so. In frontier-scale training, that's the only explanation that matters.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.