Synthetic Oracle Feedback — reveals a non human instance of → Human-Feedback Implicit Reward Shaping

explored within the theme Hiding the reward: experimental design for preference-only learning

The shaping phenomenon is usually told as a story about human judgment specifically, that humans reward 'standing upright' or 'passing cars' better than a hand-written bonus does, but the paper reports the same shaping advantage appearing with no human in the loop at all: at 1400 labels, training against the synthetic oracle's comparisons 'performs slightly better than if it had simply been given the true reward,' which the authors attribute to the learned reward function being 'slightly better shaped' than the true reward it approximates (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 7). Since the oracle's answers are generated mechanically from the exact same true reward function being compared against, this instance of shaping cannot come from a rater's independent judgment; it has to come from the comparison-based reward-fitting procedure itself, which the authors describe as assigning 'positive rewards to all behaviors that are typically followed by high reward,' a smoothing effect the raw reward function does not get. That finding complicates the 'implicit human reward shaping' framing elsewhere in the paper: at least part of the shaping benefit is a property of learning reward from comparisons at all, not a property of what a human specifically notices or cares about.