Compute vs. Human Feedback Cost Analysis — caps the marginal value of → Human-Feedback Sample Efficiency

explored within the theme The economics of human feedback

Section 4 pairs the sample-efficiency headline with an argument for why squeezing it further would not much matter in practice: right after reporting the roughly three-orders-of-magnitude reduction in interaction complexity, the discussion adds that the paper is 'already hitting diminishing returns on further sample-complexity improvements because the cost of compute is already comparable to the cost of non-expert feedback' (deep-rl-human-prefs, §"4 Discussion and Conclusions", p. 10). The footnote behind that claim puts numbers on it: roughly $25 of cloud compute for a day of Atari training against roughly $36 of minimum-wage labor for the 5,000 labels behind it, an order-of-magnitude-comparable pair of costs (deep-rl-human-prefs, §"4 Discussion and Conclusions", p. 11). That comparison reframes what the sample-efficiency result is for: the point was never to make human feedback free, only cheap enough that compute, not labor, becomes the binding constraint, and the cost analysis is the evidence that this crossover had already been reached, capping how much further gains in labels-per-task would actually be worth pursuing.