Jailbreaks — motivates → CAA as an adversarial red-teaming tool (proposed application)
The red-teaming proposal needs a reason to exist: if RLHF and finetuning reliably closed off unwanted behavior, there would be nothing further to elicit. Jailbreaks are the paper's evidence that they don't. Users routinely find inputs that make aligned, RLHF-trained chat models "output harmful content" despite that training (contrastive-activation-addition, §"9.1 Suggested future work", p. 9), which the paper treats as proof that alignment training suppresses rather than eliminates unwanted behaviors -- they remain latent, reachable by some input, just not one anyone has found yet through ordinary use. What jailbreaks don't provide is a method: they are discovered incidentally, by users probing a deployed model, not produced by any systematic process. caa-red-teaming-application is pitched directly at that gap -- not at the existence of unwanted behavior, which jailbreaks already establish, but at the difficulty of finding where it hides. The paper's own framing makes the dependency explicit, raising jailbreaks in the same paragraph as the proposal meant to address the problem they expose.