One paper's central result, traced edge by edge
Reward harmlessness without ever penalizing refusal, and a model learns to dodge instead of engage; that is how evasiveness crept into Constitutional AI's own results before the paper caught it and redesigned what crowdworkers were rewarding for.
From that redesign the page steps back to trace the whole arc: the tension curdling into evasiveness in the first place, the ordinary-RLHF baseline used for comparison, and RL-CAI's rewritten principles pulling the result above that baseline rather than just along it. One caveat closes things out: the scoring method itself was calibrated to punish evasive answers.
Constitutional AI's headline claim has a specific empirical arc that these edges trace in sequence. The helpfulness-harmlessness tension collapses into evasiveness whenever crowdworkers reward harmlessness without penalizing refusal. The helpful RLHF model traces the tradeoff curve with the HH RLHF model, the standard-RLHF reference line the paper plots against. RL-CAI eliminates evasiveness by redesigning the comparison principles so the feedback model never rewards refusal for its own sake, and RL-CAI produces a Pareto improvement, landing above the RLHF tradeoff line rather than merely moving along it. None of this is metric-neutral: the Elo score was itself tuned to penalize evasiveness, a methodological choice the paper admits compresses the harmlessness gap relative to prior work. The five edges together show a result built as much from redesigning what gets rewarded and how it gets scored as from any change to the RL algorithm.