Superposition and the case against the neuron basis
Ask which neuron represents a concept in a language model and you may be asking the wrong kind of question. A single unit can fire for several unrelated ideas at once, and this theme argues that isn't a training accident but a structural one.
It traces polysemanticity back to superposition, the idea that models represent more features than they have dimensions by packing them into overlapping, non-orthogonal directions, and shows the same fact twice: theoretically, in why the residual stream has no privileged coordinate system except for a few outlier exceptions, and empirically, in neurons and coordinates scoring worse than learned features on interpretability.
This theme's claim is that neurons, and basis coordinates generally, are the wrong unit of analysis for reading a language model, and the paper demonstrates this twice over. The premise is theoretical: polysemanticity, individual units activating for unrelated concepts, is hypothesized to arise from superposition, models packing more features than they have dimensions into an overcomplete set of non-orthogonal directions. The residual stream, the transformer's running sum of layer outputs and the paper's main object of study, is not expected to have a privileged basis, no coordinate system more inherently meaningful than any rotation of it, though the paper documents outlier dimensions, specific coordinates carrying disproportionately large values, as a partial exception. The empirical half of the same fact follows: the default basis baseline, scoring individual neurons or residual-stream coordinates directly, is shown polysemantic exactly where the corresponding learned feature is monosemantic, activating for one coherent concept. Together these findings back mechanistic interpretability's premise that meaningful analysis units must be discovered, not assumed to be neurons or coordinates.