RL-CAI (RLAIF-trained Constitutional AI model) — produces → Pareto Improvement (helpfulness vs. harmlessness)
Pareto improvement is not a property CAI asserts abstractly; it is read directly off where rl-cai’s Elo points fall in Figure 2 relative to the line traced by helpful-rlhf-model and hh-rlhf-model. Ordinary RLHF only lets you trade helpfulness for harmlessness by moving along that existing line — training harder for harmlessness pushes you further into evasiveness. RL-CAI instead lands above the line: at matched or better helpfulness, it scores more harmless than either RLHF baseline. The mechanism is supply, not cleverness in the RL algorithm itself — replacing scarce, expensive human harmfulness comparisons with abundant, principle-generated AI comparisons lets far more harmlessness training happen without re-importing the evasiveness bias baked into the earlier human-labeled HH data. (constitutional-ai, §"1.1 Motivations", p. 3)