Steering at targeted token positions (proposed extension) — would relax the quality ceiling of → Contrastive Activation Addition (CAA)

explored within the theme A behavior becomes a direction in activation space

The proposal to steer only targeted token positions is a direct response to a ceiling CAA already ran into empirically. In §4.2, the authors report exploring "a wider range of multipliers" before settling on a restricted range, because larger multipliers degraded open-ended generation quality by both GPT-4's and human readers' judgment (contrastive-activation-addition, §"4.2 Open-ended generation", p. 4). The Discussion traces that ceiling to where the vector is applied: because it is added "at every token position after the user's prompt," any single multiplier perturbs the entire continuation uniformly, which "results in a cap on the amount by which we can perturb the representations before degrading text quality" (contrastive-activation-addition, §"9.1 Suggested future work", p. 9). Targeting a smaller, more selective subset of positions is proposed as a way to relax that cap — spending the same total perturbation budget more efficiently by concentrating it where it changes the behavior rather than diffusing it across tokens where the behavior isn't relevant, potentially achieving "a better trade-off between intervention size and effect size" than blanket, all-position steering allows.