Contractor Preference-Labeling Protocol — becomes a screened operation in → Labeler screening and selection process
The 2017 paper's entire labeling workforce is recruited by having contractors sign up for a slot in a spreadsheet and then compare clips through a keyboard-driven interface, with no formal test of whether any given contractor is any good at the task (deep-rl-human-prefs, §"B Instructions Provided to Contractors", p. 15). Nothing in the 2017 pipeline screens for label quality before a contractor's judgments enter the preference database; the paper's own results note that irregular learning progress on the Hopper task is attributable to 'one contractor deviating from the typical labeling schedule,' a quality problem the pipeline had no mechanism to catch before it affected training (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 7). InstructGPT's labeler screening and selection process is what a labeling operation looks like once that gap is closed: prospective labelers are scored on four separate criteria — agreement on sensitive-speech flagging, agreement on rankings, demonstration-writing quality, and self-assessed ability to judge sensitive content across groups — before the roughly 40 who pass are allowed to contribute any training data at all (instructgpt, §"B.1 Labeler selection", p. 36). Read backward from 2022, the 2017 sign-up sheet is a labeling operation with the entire screening stage missing, and the contractor-deviation anecdote shows that gap was already visibly a problem five years before anyone built a fix for it.