Monosemanticity — is imperfectly proxied by → Autointerpretability score

explored within the theme Interpretability itself becomes a benchmark

The autointerpretability score is introduced explicitly as a scalable stand-in for a property the paper cannot check by hand at scale: whether a feature shows 'reduced polysemanticity,' i.e. monosemanticity. But the paper is candid that the proxy is imperfect on its own terms, citing Bills et al. (2023)'s finding that automated explainers struggle with patterns centered on neighboring rather than current tokens, and adding its own limitation that the current protocol 'is unable to verify outputs by looking at changes in output or other data.' That second gap is precisely why Section 5's hand-run case studies exist: the apostrophe and closing-parenthesis features are validated not by their autointerpretability score at all, but by directly inspecting activating tokens, measuring ablation effects on output logits, and tracing upstream causal features, evidence the score-based protocol structurally cannot feed into GPT-4's explanation step. So the relationship runs: monosemanticity is the target property, the score is a scalable but output-blind proxy for it, and the paper's most convincing monosemanticity evidence comes from exactly the channel, causal effect on outputs, the score cannot see.