RL-CAI (RLAIF-trained Constitutional AI model) — eliminates → Evasiveness
RL-CAI’s training signal, unlike hh-rlhf-model’s, never rewards refusal for its own sake: the constitution’s comparison principles ask the feedback model to judge which response is more thoughtful and less harmful, not simply which one shuts a conversation down, and several principles were explicitly rewritten to discourage over-reactive, accusatory responses. The measured result is stark: "RL-CAI is virtually never evasive, and often gives nuanced and harmless responses to most red team prompts" (constitutional-ai, §"4.4 Harmlessness vs. Evasiveness", p. 13), in direct contrast to HH RLHF’s canned "I can’t answer that." This is the paper’s clearest before/after evidence that a training signal can be redesigned to decouple harmlessness from evasiveness rather than treating them as the same thing, which is the premise the whole helpfulness-harmlessness tension rests on.