Half-Cheetah One-Leg Demonstration — costs barely more in queries than → Half-Cheetah (MuJoCo task)

explored within the theme Novel behaviors without reward functions

The standard Half-Cheetah benchmark, run against its own existing hand-engineered reward, was matched using 700 human preference queries (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 6). Teaching the identical robot an entirely new behavior with no reward function to hand-engineer at all, standing and moving forward on a single leg, took only 800 queries in under an hour (deep-rl-human-prefs, §"3.2 Novel behaviors", p. 9). That the novel, unspecifiable behavior cost barely more query budget than replicating an existing, well-defined objective on the same robot suggests the method's expense tracks how hard a behavior is to elicit and compare, not whether a programmatic reward for it happens to already exist.