Building the human side of the RLHF pipeline — exposes the noisy consensus beneath → InstructGPT's three-step training pipeline
The pipeline theme describes reward-model training as fitting a pairwise comparison loss to which of two outputs human labelers would prefer, language that reads as though labeler preference were a stable ground truth to be recovered. The data-infrastructure theme's own concepts complicate that reading without the pipeline narrative registering it: independent labelers agree with each other only about 73% of the time (instructgpt, §3.4 Human data collection, p. 8), and reward models cross-validated across labeler groups reach roughly 69.6% held-out accuracy versus 72.4% in-group (instructgpt, §E.2 Reward model generalization across sets of labelers, p. 52). What steps two and three of the pipeline actually optimize against, in other words, is a moderately noisy, only partly generalizing consensus among roughly 40 screened contractors, not the singular target the pipeline theme's comparison-loss description suggests.