Elicitation as audit: steering for red-teaming — inverts the control mechanism of → A behavior becomes a direction in activation space
Steering's own control machinery, a vector times a signed multiplier, is what the red-teaming proposal repurposes unchanged, only run toward the behavior instead of away from it. Nothing in how the vector is built or applied differs: the same Equation-1 average-difference vector, the same injection at every token position after the prompt, that suppresses a behavior at multiplier −1 elicits it at multiplier +1 just as reliably (contrastive-activation-addition, §"3 Method", p. 3). The paper's own framing of this inversion is explicit: "CAA could be used as an adversarial intervention to trigger unwanted behaviors in models more efficiently," on the premise that "if a behavior can be easily triggered through techniques such as CAA, it may also occur in deployment" (contrastive-activation-addition, §"9.1 Suggested future work", p. 9). What was a dial for controlling output becomes, at the same setting, a probe for testing whether safety training actually removed a behavior or just suppressed its default expression.