Synthetic Oracle Feedback — is only realizable within → Quantitative vs. Qualitative Evaluation Methodology

explored within the theme Hiding the reward: experimental design for preference-only learning

Synthetic oracle feedback only makes sense inside the quantitative half of the paper's evaluation methodology, since manufacturing the oracle's answer requires a withheld ground-truth reward function to consult, while the qualitative tasks (backflips, one-legged running) are defined precisely by having no such reward function at all (deep-rl-human-prefs, §"2.1 Setting and Goal", p. 4). Within that quantitative track the oracle does more than stand in for a human rater: sweeping its label count (350, 700, 1400, and further for Atari) isolates how the reward-learning method itself scales with feedback volume, independent of the noise or inconsistency a real rater introduces (deep-rl-human-prefs, §"3.1 Reinforcement Learning Tasks with Unobserved Rewards", p. 6). A subtlety the quantitative setup has to absorb is that a sparse true reward makes many oracle comparisons genuinely tied: in Atari, two clips often both score zero true reward, so the oracle frequently outputs indifference rather than a preference, and the method has to stay informative even when a large share of its ground-truth labels are ties (deep-rl-human-prefs, §"3.1 Reinforcement Learning Tasks with Unobserved Rewards", p. 6).