Token-level cosine similarity analysis — reveals as a feature detector → Steering vector

explored within the theme Steering doubles as interpretability

A steering vector is defined as a control mechanism -- something added to activations to shift the output distribution. Token-level cosine similarity analysis shows the same object also works in reverse, as a passive detector of the behavior it was built to induce. Taking the dot product between a steering vector and each token's residual-stream activation during an ordinary forward pass -- no addition, just measurement -- produces a signal that tracks how strongly the behavior is "present" at that token: on the refusal vector, the phrases "I cannot help" and "I strongly advise against" score positively, while "hack into your friend's Instagram account" scores negatively (contrastive-activation-addition, §"8.1 Similarity between steering vectors and per-token activations", p. 7). The same held for a myopia vector, where tokens describing choosing a reward now scored oppositely from tokens describing waiting for a larger one later (contrastive-activation-addition, §"8.1 Similarity between steering vectors and per-token activations", p. 8). A vector extracted to intervene on a behavior turns out to already encode enough about that behavior to recognize it token by token, with no intervention at all.