Labeler screening and selection process — creates the question answered by → Held-out labeler generalization test

explored within the theme Building the human side of the RLHF pipeline

Screening is what makes generalization worth testing. Because every training label came from roughly 40 contractors filtered through four criteria — including agreement with the researchers' own judgments (instructgpt, §"B.1 Labeler selection", p. 36) — the headline preference results could in principle reflect the tastes of a small, curated, researcher-aligned group rather than anything broader. The held-out experiment is built to break that dependence: the extra labelers come from the same vendors but "do not undergo a screening test" (instructgpt, §"3.4 Human data collection", p. 8), so passing the filter cannot explain their judgments. They prefer InstructGPT to GPT-3 at roughly the training labelers' rates, and five-fold cross-validated reward models predict held-out labeler groups' preferences at 69.6% versus 72.4% within-group (instructgpt, §"E.2 Reward model generalization across sets of labelers", p. 52) — a measurable but small overfitting-to-labelers cost. The test answers the narrow question the screening created; whether 40 contractors represent the broader population of users is a different and harder question.