Refusal (as a steerable behavior) — quantifies → Helpfulness-Harmlessness Tension
Constitutional AI observes a general tradeoff: increasing a model's willingness to help tends to increase its potential for harm, and increasing harmlessness training tends to make it more evasive and less helpful, so CAI frames its own contribution as shifting the Pareto frontier of that tradeoff rather than picking a point on it (constitutional-ai, §"1.1 Motivations", p. 2). That framing leaves the tradeoff as a qualitative property of a training method. CAA's Refusal behavior turns it into a signed, continuous quantity: a single steering vector, extracted from GPT-4-authored contrast pairs, that can be added or subtracted at inference time to move any already-trained chat model along that same axis, with the multiplier controlling how far (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 3). Where Constitutional AI can only compare models trained different ways and observe where each one landed on the frontier, CAA takes one fixed model and slides it along the frontier after training, at inference time, without retraining anything -- evidence that the tension is not baked irreversibly into weights but lives, at least partly, in a direction the network already represents and that can be read out and reapplied.