The economics of human feedback — supplies the cost argument behind → Building the human side of the RLHF pipeline
This theme's closing arithmetic, roughly $25 of Atari compute already comparable to the roughly $36 of minimum-wage labor behind 5,000 labels (deep-rl-human-prefs, §"4 Discussion and Conclusions", p. 11), is the paper's own case that oversight of deep RL had, for the first time, become economically practical. That argument was made at the scale of a handful of contractors who signed up for open time slots off a shared spreadsheet, briefed with a one-to-two sentence task description (deep-rl-human-prefs, §"3.1 Reinforcement Learning Tasks with Unobserved Rewards", p. 6). InstructGPT's human-data infrastructure is what building on that argument at industrial scale actually requires: not a spreadsheet but a four-criterion labeler screening process selecting roughly 40 contractors (instructgpt, §"B.1 Labeler selection", p. 36), and not thousands of labels but the roughly 33,000-prompt RM dataset alone (instructgpt, §"3.2 Dataset", p. 7). Nothing in the 2017 accounting anticipates a purpose-built web labeling interface or a fixed metadata annotation taxonomy for failures like hallucination or denigration of a protected class; those exist because the budget stopped being a few contractors on a spreadsheet and became a standing operation with its own overhead to manage. The relation between the two themes is the argument becoming the org chart: the same minutes-of-attention logic that made $25 of compute look cheap next to $36 of labor is what, scaled up, justifies building screening pipelines and interfaces rather than simply paying more contractors to do more of the same simple task.