Labeler screening and selection process
Training language models to follow instructions with human feedback — introduced
A four-criterion screening procedure (agreement on sensitive-speech flagging, agreement on rankings, sensitive demonstration writing quality, and self-assessed ability to identify sensitive content across groups) used to select the roughly 40 contractors who produced the paper's training data.
Before InstructGPT could collect a single training label, someone had to decide who was trusted to judge sensitive content in the first place. The process that narrowed a larger candidate pool down to about forty contractors is where that choice actually got made.
Four criteria narrowed the field, and the page works through each one before turning to what clearing that bar does and doesn't guarantee about data quality, closing on the open question the held-out generalization test later picks up.
The four criteria
InstructGPT's roughly 40 training labelers, hired through Upwork and Scale AI, were selected from a larger candidate pool by a four-criterion screening process designed to find people "sensitive to the preferences of different demographic groups" and "good at identifying outputs that were potentially harmful" (instructgpt, §"3.4 Human data collection", p. 7; §"B.1 Labeler selection", p. 36). The criteria were: agreement with researcher labels on flagging sensitive speech; agreement with researcher rankings of model completions by quality; a 1-7 Likert-scored "demonstration score" on a small set of sensitive prompts requiring nuanced responses; and a self-assessed account of which topics or cultural groups a candidate felt comfortable identifying sensitive content for, since demographic hiring criteria were legally unavailable. Soft cutoffs of 75% agreement on the first two criteria and a 6/7 demonstration score were applied, with the fourth, subjective criterion folded in by judgment (instructgpt, §"B.1 Labeler selection", p. 36).
What passing the filter does and doesn't buy
The Inter-labeler agreement rate — was both filter and audit for → Labeler screening and selection process edge surfaces a counterintuitive result: the screened training labelers agree with each other only 72.6 ± 1.5% of the time, while unscreened held-out labelers, sourced from the same vendors but never subjected to this test, agree among themselves at a higher 77.3 ± 1.3% (instructgpt, §"3.4 Human data collection", p. 8). Screening, in other words, did not select for homogeneity of taste; its criteria targeted sensitivity to harmful content across many demographic groups, a different property than agreement rate measures.
Creating the question that held-out generalization answers
Because every training label passed through this filter, including agreement with the researchers' own judgments, the paper's headline preference results could in principle reflect the tastes of a small, curated, researcher-aligned group rather than any broader population, a concern the Labeler screening and selection process — creates the question answered by → Held-out labeler generalization test edge names directly. The Held-out labeler generalization test test is built to probe exactly this: it recruits labelers from the same vendor pool who do not undergo the screening test, and finds they prefer InstructGPT to GPT-3 at roughly the training labelers' rate. This concept anchors the Building the human side of the RLHF pipeline theme's account of how RLHF's algorithmic story depends on an entire human-selection apparatus operating beneath it.