LLM System Design Interview, issue 29, Apr 19, 2026
The Compute-Without-Data Trap
Principal AI Engineer interview at Meta, and the interviewer asks:
“We just secured a cluster of 100,000 H100s, but we’ve completely run out of high-quality internet text. How does entering this ‘data-constrained’ regime completely invert our standard assumptions about epochs and architecture design?”
Don’t say: “Just train on the existing data for more epochs and use synthetic data.”
Why 100,000 GPUs can still stall progress, and the shift from throughput optimization to sample efficiency.
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.