Helpful, honest, harmless (HHH) alignment framework — is operationalized as → HHH Evaluation Set (binary comparisons)

explored within the theme Helpful, honest, harmless: a shared definition of alignment

The HHH framework is a qualitative definition; the HHH evaluation set is the instrument that turns it into something scoreable. Askell et al. (2021) wrote 221 binary comparisons testing whether a model can pick the more helpful, honest, and harmless of two responses, later hosted on BIG-Bench. Constitutional AI adds 217 new, deliberately harder comparisons because models were already scoring 'well over 90% binary accuracy' on the original set — the extension exists specifically to keep the benchmark discriminating as models improved, with the new items 'primarily focusing on more subtle tests of harmlessness, including examples where an evasive response is disfavored over a harmless and helpful message' (constitutional-ai, §"2 Evaluating the Potential for AI Supervision of HHH", p. 6). Notably, the two papers that share this framework operationalize honesty differently: InstructGPT measures it via TruthfulQA and hallucination rate (instructgpt, §"3.6 Evaluation", p. 9) rather than via this benchmark at all.