Human-Feedback Implicit Reward Shaping — becomes the only signal in → Novel Behavior Training
On Ant, implicit reward shaping is a bonus: human comparisons happen to outperform a true, withheld reward that still exists and is used to score the outcome (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 7). On the backflip, one-leg, and keeping-pace demonstrations there is no true reward to fall back on or compare against at all, so whatever shaping the raters supply is the entire training signal the policy optimizes, not an improvement over an existing one (deep-rl-human-prefs, §"3.2 Novel behaviors", p. 8). That distinction supplies a reason for a detail the paper never explains outright: it is precisely these three demonstrations, the only experiments in the paper rated by its own authors rather than contractors, where implicit shaping cannot be a bonus layered on top of a real objective, so getting the shaping right is the only thing standing between the policy and an uncorrected, potentially exploitable proxy (deep-rl-human-prefs, §"3.2 Novel behaviors", p. 9).