Inter-labeler agreement rate — caps the resolution of → Win rate (preference against baseline)

grounded in Training language models to follow instructions with human feedback · explored within the theme Automatic overlap scores versus human and preference judgments

A win rate is a proportion of human choices, so it can never be more precise than the humans making those choices are consistent. InstructGPT's training labelers agree with one another 72.6 ± 1.5% of the time, and held-out labelers 77.3 ± 1.3% (instructgpt, §"3.4 Human data collection", p. 8), meaning roughly a quarter of individual comparisons are effectively coin flips. This sets an interpretive floor under every win-rate figure in the paper: the headline result — 175B InstructGPT preferred to GPT-3 85% of the time (instructgpt, §"4.1 Results on the API distribution", p. 11) — towers over the noise, but small gaps between adjacent models, such as PPO versus PPO-ptx, sit inside it and should not be read as rankings. Agreement also bounds the ceiling: a model cannot be preferred more consistently than evaluators agree on what "better" means. The paper leans on the comparison to Stiennon et al. (2020)'s 73 ± 4% researcher-researcher agreement to argue that its much broader, more subjective task distribution did not degrade label quality below usability.