LLM Agents Interview Questions, issue 15, Mar 9, 2026

The SWE-Bench Proxy Trap

AI Engineer interview at Google DeepMind, and the interviewer asks:

Your new autonomous coding agent is hitting 40% on SWE-Bench. The PRs pass all historical unit tests perfectly. But when you manually review the code, the patches are technically incorrect and introduce massive regressions. Why is your agent passing the test but failing the engineering task?

Don’t say: The LLM is hallucinating code syntax, or we need to increase the context window so it sees more of the codebase.

Agents scoring well on repository benchmarks often exploit underspecified tests rather than solving the underlying engineering problem.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.