The economics of human feedback — prices the implicit budget of → Scaling oversight by predicting reward on unlabeled episodes

hindsight · grounded in Deep Reinforcement Learning from Human Preferences · part of the supertheme Scalable oversight made real: from 2016 proposals to InstructGPT's RLHF pipeline

None of this theme's four proposals, supervised reward learning, active reward learning, unsupervised value iteration, unsupervised model learning, mentions cost, dollars, or contractor wages, but each is at bottom a proposal about rationing a scarce resource: an agent that observes its true reward on only a small fraction of timesteps yet is evaluated throughout (concrete-problems, §"5 Scalable Oversight", p. 11). The 2017 paper takes exactly that scarcity framing and prices it: its stated budget is minutes of a non-expert human's attention, and its headline result is learning from feedback on "less than 1% of our agent's interactions with the environment" (deep-rl-human-prefs, §Abstract, p. 1). The paper's own closing cost accounting, roughly $25 of compute against roughly $36 of minimum-wage labor for 5,000 labels (deep-rl-human-prefs, §"4 Discussion and Conclusions", p. 11), is the number the 2016 theme needed but never had: evidence that the small fraction its semi-supervised framework assumed could be afforded actually was affordable, at least at this scale. Read this way, the theme's four mechanisms are not simply different technical routes to the same statistical goal; they are variations on a single budget request, and the 2017 paper is the first to itemize what that request would actually cost and pay.