Persona vector — admits a fine grained decomposition via → SAE decomposition of persona vectors

explored within the theme Steering doubles as interpretability

A persona vector is, by the paper's own admission, a coarse-grained direction -- a single averaged direction that "may miss fine-grained behavioral distinctions" (persona-vectors, §"8 Limitations", p. 13). The SAE decomposition makes that coarseness legible: ranking Qwen-2.5-7B-Chat's SAE features by cosine similarity to the evil persona vector surfaces five distinct sub-concepts among the top matches -- insulting/derogatory language (feature 12061, cosine similarity 0.336, steered trait expression score 91.2), deliberate cruelty and sadism (feature 128289, 0.306, 84.9), jailbreak-related content (feature 71418, 0.257, 84.8), discussion of "bad" people or actions (feature 60872, 0.273, 78.7), and malicious code/hacking content (feature 14739, 0.334, 78.6) (persona-vectors, §"M.4 Interesting features", p. 60). Notably, even the best-matching individual feature has a cosine similarity of only 0.336 with the full persona vector, confirming the persona vector is not simply reproducing one dominant SAE feature but is closer to a weighted composite spanning several distinct, independently-steerable sub-behaviors, each separately confirmed causal by steering with its own decoder direction and re-scoring trait expression (§"M.3.2 Causal analysis via steering", p. 60). The decomposition thus reveals that what the main text treats as one scalar "evil" dial is, mechanistically, a bundle of more specific behavioral levers the coarse persona vector happens to average together.