Contrastive Activation Addition (CAA) — repurposes for steering relative to → Representation Engineering (Zou et al. 2023)

explored within the theme The activation engineering family: from single pairs to datasets

Zou et al.'s Representation Engineering already tested Mean Difference as one way to locate and extract representations of concepts like honesty and emotion, so CAA is not introducing MD to this lineage — it is repurposing it. Two changes mark the repurposing: contrast-pair quality and goal. RepE's contrastive prompts are less tightly matched, whereas "CAA employs an optimized multiple-choice format that results in more closely paired contrastive prompts that differ by only a single token" (contrastive-activation-addition, §"2 Related work", p. 2), which sharpens the extracted direction by removing incidental prompt variation. More fundamentally, RepE's project is representation extraction — characterizing what a model represents — while CAA "build[s] on this work by focusing on steering rather than representation extraction, experimenting with a broader range of behaviors, and comparing steering to system-prompting and supervised finetuning" (contrastive-activation-addition, §"2 Related work", p. 2). So where RepE treats the mean-difference direction as a diagnostic to be read, CAA treats the same kind of direction as a lever to be pulled, with the tightened contrast format making that lever more precise.