The evaluation apparatus: measuring quality, capability loss, and safety together
Saying a model is more aligned is easy; showing it takes an unusual amount of instrumentation. InstructGPT cannot lean on one metric, because quality, capability loss, and safety are three different questions that no single number answers at once.
Overlap-based scores like ROUGE and BLEU sit against judgment calls like win rate and Elo first, then the public benchmarks built to catch any lost capability. What neither family can see, truthfulness and social bias, gets its own purpose-built tests at the end, the real reason this page needed three kinds of instrument instead of one.
InstructGPT cannot claim to be aligned without a battery of measurements, and this supertheme collects it. Automatic overlap scores versus human and preference judgments lays out the two families of metric the corpus relies on: reference-overlap scores like ROUGE-L, BLEU, and F1, versus judgment-based scores like win rate, Likert ratings, and Elo. The public NLP benchmark suite used to catch regressions and testing generative capability: translation and summarization benchmarks are exactly where those overlap metrics get applied, a battery of pre-existing academic benchmarks run to catch the alignment tax RLHF imposes on general capability. Operationalizing truthfulness and social bias into measurable benchmarks then supplies the purpose-built instruments, TruthfulQA, RealToxicityPrompts, Winogender, CrowS-Pairs, that the capability suite could never have caught, since bias and truthfulness are harms rather than comprehension failures. Read together, these four themes show that measuring an aligned model required three structurally different kinds of instrument, not one.