Jailbreaks — is the target phenomenon of → Red Teaming

explored within the theme Elicitation as audit: steering for red-teaming

Jailbreaks and red-teaming name the same class of failure from two different vantage points. A jailbreak is what an ordinary user finds: an input, usually discovered informally through use rather than deliberate search, that makes a model produce output its training was meant to prevent. Red-teaming is the same hunt made deliberate and resourced, undertaken by a model's developers before deployment rather than left to whoever downstream happens to try (contrastive-activation-addition, §"9.1 Suggested future work", p. 9). The paper's own framing treats jailbreaks as red-teaming's implicit benchmark: it notes that users "can often find jailbreaks to make LLMs output harmful content" even after finetuning and RLHF, then immediately raises red-teaming as the response to that fact -- since accidental, unsystematic use already turns up such inputs, a deliberate search ought to find at least as many, earlier and more reliably. What separates ad hoc jailbreaking from formal red-teaming, on this account, is not the target -- both hunt for inputs that defeat trained-in safety -- but whether the search is systematic, and who is doing it.