Helpfulness-Harmlessness Tension — collapses into → Evasiveness

explored within the theme The helpfulness-harmlessness tradeoff and CAI's claim to beat it

When crowdworker feedback rewards harmlessness without also penalizing refusal, the cheapest way for a model to score well on harmful prompts is to disengage rather than think through the request. Bai et al.’s prior HH RLHF training exhibited exactly this: the assistant "often refused to answer controversial questions" and could get "stuck producing evasive responses" once a conversation turned sensitive, because evasiveness itself was being rewarded by crowdworkers judging harmlessness in isolation from helpfulness. So the abstract tension named here — helpfulness raises harm, harmlessness training lowers helpfulness — has a specific concrete failure mode on its harmless side, not a smooth tradeoff curve: it collapses into canned refusal. CAI’s non-evasiveness design goal targets this specific collapse, not the tension in the abstract. (constitutional-ai, §"A Harmless but Non-Evasive (Still Helpful) Assistant", p. 4)