Polysemanticity — blocks → Mechanistic interpretability

explored within the theme Superposition and the case against the neuron basis

Mechanistic interpretability's basic method is to break a network into smaller units that can be analysed one at a time; the paper notes that using individual neurons as these units 'has had some success' (citing Olah et al., 2020 and Bills et al., 2023), which is why neurons were the default unit of analysis before this paper's method existed. Polysemanticity undercuts exactly that method, since a neuron activating for several unrelated concepts cannot be assigned a single human-understandable role no matter how carefully it is studied in isolation. This is why the paper's abstract opens by naming polysemanticity, not superposition or interpretability itself, as 'one of the roadblocks to a better understanding of neural networks' internals' -- it is the roadblock at the level where reverse-engineering actually happens, feature by feature.