Steering doubles as interpretability — surfaces artifacts explained by → Superposition and the case against the neuron basis
CAA's activation-visualization technique already demonstrates, on a much larger and RLHF-tuned model than this paper studies, the same structural claim this theme argues from first principles: projecting Llama 2 7B/13B Chat's residual-stream activations by PCA on a contrastive dataset shows "behavioral clustering" -- separation by target behavior rather than by superficial prompt format -- emerging suddenly "around one-third of the way through the layers" (contrastive-activation-addition, §"3.2 Visualizing activations for contrastive dataset analysis", p. 4), evidence that meaningful structure lives in directions of activation space rather than in any privileged coordinate, arrived at independently of anything this paper's own case studies show on Pythia. What this theme adds back is a mechanistic account the steering side has no access to: when this paper searches by hand for the residual-stream dimension nearest the apostrophe feature, the top-activating dimensions it finds are not generic coordinates but specifically outlier dimensions, disproportionately large-magnitude coordinates the paper elsewhere attributes to "the Adam optimiser saving gradients with finite precision in the residual basis" (sparse-autoencoders, §"D.3 Feature Search Details" and §"C.2", pp. 15-17) -- a specific, named cause for why some directions in an otherwise unprivileged basis behave as if they were privileged after all. A steering paper's PCA-based behavioral clustering and this paper's dictionary-based feature search independently corroborate that directions, not neurons, carry the structure; only this theme explains why a handful of those directions look disproportionately important in raw coordinate space to begin with.