Issue 16, Oct 5, 2026
The Side-Effect Blindspot
Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
“Your computer-use agent gets the task: ‘Add a blue mug under $20 to the cart, but don’t check out.’ Your outcome check confirms the mug is in the cart and passes the run. What did you fail to measure, and how do you enforce the ‘don’t’ part in evals and in production?”
Don’t say: “Add Do NOT check out to the system prompt and verify at the end.”
Why your computer-use agent can pass every eval check while silently draining a user's credit card, and how to enforce negative constraints outside the LLM.
Share this trapShare on LinkedIn
The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.