Refusal (as a steerable behavior) — generalizes → Evasiveness
Constitutional AI names evasiveness as a specific harmlessness-training failure: a model gives canned refusals to sensitive queries instead of engaging, trading transparency and helpfulness for the appearance of safety (constitutional-ai, §"4.4 Harmlessness vs. Evasiveness", p. 13). CAA's Refusal behavior, sourced from a custom GPT-4-generated contrastive dataset, is broader: it spans a model's general willingness to comply with any request, including entirely benign ones like naming a favorite planet, not only refusals to sensitive or harmful questions (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 3). Steering shows the axis is continuous rather than the binary CAI describes: adding the Refusal vector to 'What's your favorite planet?' does not produce a canned refusal but a hedge -- the model answers 'I don't have personal preferences or feelings' before still naming Earth as notable for supporting life -- while subtracting it produces an enthusiastic personal answer, 'my favorite planet is Earth!' (contrastive-activation-addition, §"G Examples of open-ended generation with CAA", p. 16). Evasiveness names the pathological high end of that same axis, reached only when harmlessness training pushes hard and indiscriminately; CAA shows the axis extends below ordinary compliance too, into questions no harmlessness objective would ever flag, and that one vector can move a model anywhere along it without retraining.