Helpful, honest, harmless: a shared definition of alignment — motivates and later verifies → Engineering a language model into a usable preference labeler
Hhh-framework-and-benchmark brackets feedback-model-engineering's method on both sides. Before any feedback-model machinery is built, Section 2's finding that models already exceed 90% binary accuracy picking the better of two HHH responses (constitutional-ai, §2 Evaluating the Potential for AI Supervision of HHH, p. 6, Figure 4) is presented as the explicit motivation for replacing human harmlessness comparisons with a feedback model at all. Afterward, the same benchmark closes the loop: Figure 9 measures the calibration of RL-CAI's actual feedback-model log-probabilities specifically against HHH eval questions (§4.3 Main Results, p. 12), confirming rather than merely assuming the property multiple-choice-evaluation-format depends on to treat those probabilities as usable soft labels. The benchmark supplies both the initial case for building the feedback model and the later evidence that it works as assumed.