Human fingerprints in the learning curves
A learning curve in the 2017 paper is not a clean record of an algorithm; it also records how fast a rater clicked, how evenly they showed up, and whether they were a contractor or an author. Response times of three to five seconds per query turn a label count into the paper's headline efficiency claim; a contractor's schedule deviation explains Hopper's irregular curve, and Half-Cheetah's feedback came from a rushed author instead.
The walk follows these fingerprints through Enduro's hard exploration case and a synthetic oracle showing the same shaping without any human involved, closing on the novel-behavior demonstrations, where that shaping is the only signal the policy gets.
The 2017 learning curves record the humans in the loop as much as the algorithm: how fast raters respond sets the headline arithmetic, uneven labeling schedules visibly change a curve, switching from contractors to authors changes the experimental condition itself, and human judgment leaks into the reward as shaping whether or not anyone intended it to. Contractor preference labeling grounds the headline efficiency claim of human feedback sample efficiency, since converting a label count into 3-to-5-second response times is what turns 700 or 5,500 labels into the well-under-an-hour and under-eight-hours figures the paper's efficiency claim actually rests on. A compute-vs-human cost analysis then caps the marginal value of squeezing that efficiency further, showing roughly $25 of cloud compute against roughly $36 of labor for the same day of Atari training, a comparable pair of costs the paper reads as diminishing returns on chasing still-fewer labels. The Enduro task supplies the hard exploration case for human feedback's implicit reward shaping: raters told only to reward progress toward passing cars give the policy a signal far richer than random exploration ever finds on its own, closing an exploration gap that neither the true reward nor synthetic oracle labels close as well. The Hopper task exposes the scheduling fragility of contractor preference labeling, its visibly irregular learning curve traced by the paper to a single contractor who deviated from the expected labeling schedule; label annealing's decaying query-rate schedule is correspondingly only approximated under contractor labeling in general, exact when a synthetic oracle supplies labels on demand but uneven whenever real contractors set their own pace. The Half-Cheetah task was exempted from contractor preference labeling entirely, its benchmark feedback supplied by an author under time pressure rather than by an unfamiliar contractor, which makes its curve a different experimental condition from the rest of the figure; the half-cheetah one-leg demonstration nonetheless costs barely more in queries than the ordinary half-cheetah task, 800 versus 700, suggesting query cost tracks how hard a behavior is to elicit rather than whether a hand-written reward for it already exists. Synthetic oracle feedback reveals a non-human instance of the same implicit reward shaping, since training against a mechanical oracle at 1,400 labels still slightly outperforms the true reward it approximates, showing that at least part of the shaping benefit comes from learning through comparisons at all, not from anything a human rater specifically notices. And human feedback's implicit reward shaping becomes the only signal in novel behavior training: the backflip, one-leg, and keeping-pace demonstrations have no true reward to fall back on, so getting the shaping right, in the very experiments rated by the paper's own authors rather than contractors, is the only thing standing between the policy and an uncorrected proxy.