Contractor Preference-Labeling Protocol — grounds the headline efficiency claim of → Human-Feedback Sample Efficiency
The claim that agents learn from feedback on 'less than 1%' of environment interactions only becomes a claim about human labor once it is converted into time, and the contractor protocol supplies exactly that conversion: contractors responded to the average query in 3-5 seconds, so the real-human experiments required between 30 minutes and 5 hours of human time in total (deep-rl-human-prefs, §"3.1 Reinforcement Learning Tasks with Unobserved Rewards", p. 6). At that rate, the 700 labels used for a typical MuJoCo task and the 5,500 labels used for a typical Atari task are not abstract counts but concrete labor budgets of well under an hour and under eight hours respectively, which is what turns the 'roughly 3 orders of magnitude' reduction in Section 4 into a claim about wall-clock human attention rather than just query count (deep-rl-human-prefs, §"4 Discussion and Conclusions", p. 10). The efficiency headline therefore depends on an empirical fact about the contractor pipeline's throughput and not only on the learning algorithm's sample-efficiency; a slower interface or less fluent raters would erode the number even with an identical count of comparisons.