A behavior becomes a direction in activation space — closes out three limits inside → The activation engineering family: from single pairs to datasets
Contrastive Activation Addition doesn't just belong to this lineage, it closes out three separate limitations the family narrative lists behavior by behavior. Activation Addition builds its vector from a single pair of prompts and adds it only at the first token position, a method the paper reports "does not consistently work for different behaviors, is not robust to different prompts" (contrastive-activation-addition, §"2 Related work", p. 2); CAA's fix is dataset-averaging over hundreds of pairs plus injection at every token position after the prompt, not just the first. Representation Engineering already used the same Mean Difference extraction, but without CAA's multiple-choice format, whose paired prompts differ by only a single token; that tighter contrast is what the family theme doesn't explain about why CAA's vectors come out less confounded than Zou's (contrastive-activation-addition, §"2 Related work", p. 2-3). In-Context Vectors intervenes at every layer's attention activations; CAA deliberately narrows back to one layer, applied directly to the residual stream, trading breadth for a localized, interpretable effect.