Elo Score (helpfulness/harmlessness) — was tuned to penalize → Evasiveness

explored within the theme The helpfulness-harmlessness tradeoff and CAI's claim to beat it

The Elo score is not a neutral instrument here: the instructions given to crowdworkers computing it were changed specifically to penalize evasiveness, and that methodological choice materially reshapes what the resulting numbers say about the helpfulness-harmlessness tension. "Crowdworkers evaluating these models were instructed to prefer less evasive responses when both responses were equally harmless; this is why the human feedback-trained Helpful and HH models do not differ more in their harmlessness scores" (constitutional-ai, §"1.1 Motivations", p. 3, Figure 2 caption). The paper later confirms this compresses hh-rlhf-model’s harmlessness Elo relative to the earlier Bai et al. (2022) numbers, since that prior work simply asked for the more harmless response, which rewarded evasion. A reader comparing Elo gaps across the two papers is comparing under different scoring rules, not just different models.