The economics of human feedback — underwrites choosing → Comparisons as the atomic unit of human feedback

part of the supertheme Deep RL from Human Preferences: the 2017 proof of concept

Every property that makes comparisons the paper's chosen feedback format is also a property the labeling budget specifically needs, a connection neither theme states on its own. A comparison costs about the same few seconds as any other judgment, contractors responded to the average query in 3-5 seconds (deep-rl-human-prefs, §"3.1 Reinforcement Learning Tasks with Unobserved Rewards", p. 6), so switching from absolute scores to comparisons does not slow the labeling pipeline down. What it buys instead is freedom from calibration: because the workforce is interchangeable contractors who sign up for scattered time slots rather than one trained rater working continuously (deep-rl-human-prefs, §"B Instructions Provided to Contractors", p. 15), an absolute-score protocol would require every contractor to hold the same internal scale, an overhead a minutes-of-attention budget has no room for. A comparison needs no shared scale at all, whichever contractor is on shift only has to say which of two clips is better, not where either sits on an absolute axis. That is the specific way the economics of human feedback, distinct from the accuracy argument the comparisons-vs-absolute-scores ablation already makes, explains the choice of comparisons: the paper's own accounting of the workforce, 5,000 labels for about $36 of minimum-wage labor (deep-rl-human-prefs, §"4 Discussion and Conclusions", p. 11), only pencils out because no part of that budget has to be spent aligning raters' scales.