LLM System Design Interview, issue 20, Nov 17, 2025

Why Raw User Data Will Always Fail Fine-Tuning

Senior ML Engineer interview at Perplexity, and the interviewer asks:

Your PM wants to fine-tune our new model on a random 1M sample of live user prompts to improve real-world performance. You tell them it’s a terrible idea. Why?

Don’t say: Because the data is noisy, has PII, and needs to be cleaned.

The goal isn’t to model the average user - it’s to model the valuable one.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.