Advanced Reinforcement Learning Interview Questions, issue 7, Feb 2, 2026

The Dynamics Invariance Trap

Senior AI Engineer interview at Boston Dynamics, and the interviewer asks:

We have 10k hours of data from a robot walking on concrete. We want to use Hindsight Relabeling (HER) to bootstrap a new policy for walking on sand. Is this a good idea?

Relabeling trajectories across environments with different contact mechanics poisons value estimates by assuming physics that never existed.

The full answer, with the mechanism and the arithmetic, is for paid subscribers on Substack.

Read it on Substack

Get the next one

Free on Substack. Unsubscribe in one click.