HHH Evaluation Set (binary comparisons) — is hosted as a subset of → BIG-Bench

explored within the theme Helpful, honest, harmless: a shared definition of alignment

The HHH evaluation set has two separate homes with two separate provenances. Its original 221 comparisons live inside BIG-Bench, the large, crowd-collaborative benchmark suite assembled from contributions across many research groups (constitutional-ai, §"B Identifying and Classifying Harmful Conversations", p. 20). Constitutional AI's own 217 additional, harder comparisons are not folded into BIG-Bench in this paper — they are distributed only through the paper's own GitHub repository. So 'the HHH eval' as used by Constitutional AI is really a composite of a small, versioned public subset hosted inside a much larger benchmark suite plus a private extension attached to one paper, a distinction that matters for anyone trying to reproduce or extend the evaluation later.