Label Annealing — is only approximated under → Contractor Preference-Labeling Protocol

explored within the theme The economics of human feedback

Label annealing is specified as a precise schedule, decaying the query rate roughly in proportion to the inverse of elapsed timesteps, and that schedule is followed exactly when labels come from the synthetic oracle, since the experiment can request labels on whatever cadence it likes (deep-rl-human-prefs, §"A Experimental Details", p. 14). Real contractors break that precision: the paper states plainly that 'when training with real human feedback, we attempt to similarly anneal the label rate, although in practice this is approximate because contractors give feedback at uneven rates' (deep-rl-human-prefs, §"A Experimental Details", p. 14). The gap between the two conditions is not cosmetic, since the whole point of annealing is to control how much weight new states get relative to old ones as training proceeds; a labeling protocol built around contractors self-scheduling time slots cannot deliver that control as tightly as a synthetic oracle can, so label annealing is best read as a clean idealization the contractor pipeline can only approximate.