Labeler screening and selection process — creates the question answered by → Held-out labeler generalization test
Screening is what makes generalization worth testing. Because every training label came from roughly 40 contractors filtered through four criteria — including agreement with the researchers' own judgments (instructgpt, §"B.1 Labeler selection", p. 36) — the headline preference results could in principle reflect the tastes of a small, curated, researcher-aligned group rather than anything broader. The held-out experiment is built to break that dependence: the extra labelers come from the same vendors but "do not undergo a screening test" (instructgpt, §"3.4 Human data collection", p. 8), so passing the filter cannot explain their judgments. They prefer InstructGPT to GPT-3 at roughly the training labelers' rates, and five-fold cross-validated reward models predict held-out labeler groups' preferences at 69.6% versus 72.4% within-group (instructgpt, §"E.2 Reward model generalization across sets of labelers", p. 52) — a measurable but small overfitting-to-labelers cost. The test answers the narrow question the screening created; whether 40 contractors represent the broader population of users is a different and harder question.