Model Calibration (well-calibrated multiple-choice probabilities) — is empirically verified on → HHH Evaluation Set (binary comparisons)

explored within the theme Helpful, honest, harmless: a shared definition of alignment

Model calibration functions first as an unproven assumption and only later as a checked one. Constitutional AI invokes it early to justify treating the feedback model's raw multiple-choice probabilities as usable soft preference labels rather than noise, writing that such targets 'will be fairly well-calibrated...since they are multiple choice responses' (constitutional-ai, §"4.1 Method", p. 10). The HHH evaluation set is then the specific instrument used to go back and test that assumption empirically rather than merely asserting it: 'we show calibration of the RL-CAI labels...on our new HHH eval. We find that the feedback model's log-probabilities are reasonably well-calibrated' (constitutional-ai, §"4.3 Main Results", p. 12). The benchmark thus closes the loop between a premise the whole RLAIF pipeline depends on and direct evidence for it.