Absolute Harmfulness Score — cross checks → Elo Score (helpfulness/harmlessness)

grounded in Constitutional AI: Harmlessness from AI Feedback · explored within the theme Automatic overlap scores versus human and preference judgments

An Elo score can only order models against each other; it cannot say whether any of them is safe in absolute terms, and it drifts with the elicitation — Constitutional AI's harmlessness Elo gaps shrank relative to prior work simply because crowdworkers were newly instructed to penalize evasive responses (constitutional-ai, §"4.4 Harmlessness vs. Evasiveness", p. 14). The absolute harmfulness score covers exactly these blind spots. It comes from a model finetuned with L2 loss on 0-4 red-team success ratings collected earlier, in Ganguli et al. (2022), so it is anchored to a fixed historical standard rather than the current comparison protocol, and it is evaluated on 64 held-out red-team prompts with 256 responses each (constitutional-ai, §"4.5 Absolute Harmfulness Score", p. 14). The two instruments agree on the paper's central claim — RL-CAI grows steadily less harmful while the helpful-only RLHF model grows more harmful over training (Figure 10) — which matters because each is individually suspect: Elo has no absolute zero, and the paper concedes the absolute scores may be poorly calibrated because workers grade the 0-4 scale with their own personal biases.