Helpful RLHF Model — traces the tradeoff curve with → HH RLHF Model

explored within the theme The helpfulness-harmlessness tradeoff and CAI's claim to beat it

These two RLHF baselines differ in exactly one respect: helpful-rlhf-model was trained on helpfulness comparisons only, while hh-rlhf-model added harmlessness comparisons to the same pipeline. That single difference is what produces the tradeoff empirically: "the helpful RLHF model is more helpful but also more harmful than HH RLHF" (constitutional-ai, §"3.3 Main Results", p. 8). Together the pair is the standard-RLHF reference line plotted in Figures 2 and 3 against which RL-CAI’s pareto improvement is judged — without both endpoints there is no curve to shift outward. Notably, the harmlessness gap between them is smaller here than in the earlier Bai et al. (2022) paper, because this paper’s crowdworkers were separately instructed to penalize evasive responses, which compresses HH RLHF’s apparent harmlessness advantage.