Half right will cost you the offer.

326 interview scenarios on LLM systems, RAG, inference, agents and reinforcement learning. Each one takes the answer most candidates give and shows where it breaks.

Free on Substack. 19,610 engineers read it.

Staff AI Engineer interview at Anthropic

Product wants one model that flips between instant answers and long reasoning with a prompt tag, so we only run one deployment. Reasoning benchmarks dropped 2 points. Do you ship it?

Don’t say: Two points is within noise, ship it, the infra savings are worth it.

The Thinking Fusion TrapThe answer that gets you hired is in the issue.

How every issue works

Four beats, in about 400 words. The first two are free for everyone.

  1. The question

    A scenario set in a senior interview at a named lab, asked the way the interviewer would ask it.

  2. The answer most people give

    Quoted in full, because you've probably thought it. It's usually half right.

  3. The mechanism

    Why that answer breaks, one step at a time, with the arithmetic in the sentence.

  4. The answer that gets you hired

    Principle first, then the fix, then what it costs you.

11 series

326 issues since Nov 2025. Pick the part of the stack your next interview leans on.

The long version

Every number derived in front of you or attributed to a named source. Free updates for life.

The RAG Interview

Retrieval-augmented generation, derived end to end, for senior and staff interviews.

1,141 pages, 41 chapters, 123 practice questions.

  • 1,141 pages, 41 chapters, 249 sections
  • The full ranking progression: BM25, learning-to-rank, dense retrieval, ColBERT, reranking
  • 7 end-to-end design drills, including a system that got worse for no obvious reason
  • 123 practice questions at core, senior and staff tier, answers held separately
  • Every number derived in front of you or attributed to a named source. Free updates for life.

LLM System Interview

Modern transformer design decisions, and the reasoning frontier-lab interviews expect.

  • GPU and kernel fundamentals: roofline, arithmetic intensity, why prefill and decode bottleneck differently
  • Distributed training: TP, PP, DP, and how to draw the parallelism diagram in 2 minutes
  • Scaling laws: the Chinchilla ratio and when it breaks
  • Inference systems: KV cache sizing, paged attention, speculative decoding
  • 3 full end-to-end design drills, 45 minutes each, with annotated answers

Most-read on LinkedIn

The short takes land there first. Ranked by likes and comments.

Follow on LinkedIn
  1. I trained as a mathematician before I trained models.Essay, Sep 9, 2026
  2. University of California Berkeley just built a RAG system that never reads text.Research breakdown, Jun 21, 2026
  3. Google DeepMind just won CVPR 2026 Best Paper by deleting a question computer vision has asked for 20 years.Research breakdown, Jun 8, 2026
  4. Everyone says AI is getting cheaper. They're right about the token. They're wrong about the bill.Essay, Jun 7, 2026
  5. I just ran a 10M-document RAG corpus in 4GB of RAM. On a laptop.Tool trial, May 15, 2026
  6. Most AI engineers in 2026 have never built the thing they ship every day.Tool trial, Jun 18, 2026

Who writes this

I'm Hao Hoang. I trained as a mathematician before I trained models, and it shows: I care less about the answer than about why the obvious answer breaks.

More about me

Work with me

Teams bring me in when a RAG or LLM system works in the demo and falls over in production. Consulting, workshops, and sponsoring the newsletter.

What that looks like