Monosemanticity — benchmarks → Dictionary feature

explored within the theme Superposition and the case against the neuron basis

The paper tests monosemanticity with an activation histogram technique that 'only works for dictionary features that activate for a small set of tokens,' and applied at that resolution it finds the standard is met narrowly rather than globally: the apostrophe feature (Figure 4) does not fire on all apostrophes, and two separate dictionary features are needed to cover apostrophes in '[I/We/They]'ll'-type contexts and in '[don/won/wouldn]'t'-type contexts respectively. So the dictionary does not resolve polysemanticity by finding one feature per human concept; it resolves it by proliferating many narrower, context-specific features that each individually pass the monosemanticity test. Meeting the bar this way still counts as success because the paper's other case study, the layer-5 closing-parenthesis feature, shows a single narrow feature can still be causally central to a task-relevant behaviour.