LLM System Design Interview, issue 20, Nov 17, 2025
Why Raw User Data Will Always Fail Fine-Tuning
Senior ML Engineer interview at Perplexity, and the interviewer asks:
“Your PM wants to fine-tune our new model on a random 1M sample of live user prompts to improve real-world performance. You tell them it’s a terrible idea. Why?”
Don’t say: “Because the data is noisy, has PII, and needs to be cleaned.”
The goal isn’t to model the average user - it’s to model the valuable one.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.