Inter-labeler agreement rate — was both filter and audit for → Labeler screening and selection process

explored within the theme Building the human side of the RLHF pipeline

Agreement statistics enter the pipeline twice, on both sides of the causal arrow. During selection, two of the four screening criteria are agreement measures — candidates had to match researcher labels on sensitive-speech flagging and on output rankings, with soft cutoffs at 75% (instructgpt, §"B.1 Labeler selection", p. 36). After training, inter-labeler agreement is reported as evidence the task is well-posed: training labelers agree with each other 72.6 ± 1.5% of the time (instructgpt, §"3.4 Human data collection", p. 8). The surprise is that the unscreened held-out labelers agree among themselves more — 77.3 ± 1.3% — than the screened group does. Screening, it turns out, did not select for homogeneity, plausibly because its criteria targeted sensitivity to harmful content across many demographic groups rather than similarity of taste. The comparison also bounds what the numbers can validate: agreement near the 73 ± 4% researcher-researcher rate of prior summarization work shows the ranking task is tractable, not that the selected group's preferences are the right ones to encode.