Automatic overlap scores versus human and preference judgments — adjudicates the baseline comparison in → Instruction following as an alternative to next-token prediction
Instruction-following-objective stakes its founding argument on beating FLAN and T0, the two prior cross-task instruction-tuning baselines, but the theme itself does not say how that contest was decided. The verdict comes from this theme's win rate: in a head-to-head comparison, 175B InstructGPT outputs were preferred over FLAN 78% of the time and over T0 79% of the time (instructgpt, §"4.1 Results on the API distribution", p. 13). The paper's diagnosis of why is itself measurement-driven: public instruction-tuning compilations emphasize tasks that are easy to score automatically, such as classification and QA, which make up only about 18% of real API traffic, while open-ended generation and brainstorming, poorly represented in FLAN and T0, make up 57%. The claim that a genuinely different objective beats fixed-compilation tuning is a claim quality-measurement-methods' win rate was built to test, not one instruction-following-objective could settle on its own terms.