The theory of the obstacle designs the tool
A tool built to fix a problem tends to carry the shape of that problem's own explanation, and this paper's method is a clean case of that: the theory of why interpretation is hard dictates, almost mechanically, what the fix has to look like.
These seven connections trace that chain from cause to cure: superposition explains polysemanticity, polysemanticity is what blocks reading a model directly, and that same diagnosis specifies an overcomplete, sparsity-penalized architecture rather than a vaguer fix, with monosemanticity set as its success criterion and the residual stream's mostly-absent privileged basis, and its documented exceptions, following as consequences of the same theory.
Seven edges trace one causal chain from diagnosis to design. Superposition manifests as polysemanticity, since representing more features than neurons forces shared, non-orthogonal directions, and polysemanticity blocks mechanistic interpretability, because a neuron playing several unrelated roles resists being assigned one. That same superposition motivates sparse autoencoders, dictating an overcomplete hidden layer and an l1 penalty rather than just gesturing at a fix, and monosemanticity benchmarks dictionary feature as the success criterion this design targets. The remaining edges are site facts entailed by the same theory: privileged basis is absent from residual stream, which licenses treating individual coordinates as no better than random directions; outlier dimensions partially restores privileged basis, a narrow, documented exception to that absence; and MLP SAE training confronts privileged basis directly, since the MLP sublayer is the paper's own worked example of a space that does have one, forcing different hyperparameters there than on the residual stream.