Apostrophe dictionary feature (feature 556) — is benchmarked against → Default (neuron/residual-stream) basis baseline
When the authors searched the residual stream's raw coordinate basis for the dimension most associated with apostrophes (Appendix D.1/D.3), it took checking ten candidates: the dimension eventually shown in Figure 11 (dimension 21) was only the 10th-highest-magnitude dimension the apostrophe feature itself reads from, since the 9 dimensions with larger weights didn't show apostrophes among their own top activations the way dimension 21 does. Both directions pick up apostrophes at their highest activation range, but Figure 11's caption notes the basis dimension 'only explains a very small fraction of the variance for middle-to-lower activation ranges,' exactly where the dictionary feature stays clean. The search also turned up a structural link between them: the apostrophe dictionary feature's own top-reading direction 'mainly reads from an outlier dimension' of the residual stream (Dettmers et al., 2022), and when the authors separately searched the default basis for its most positive and most negative apostrophe-related dimensions, three of the four extremes checked were confirmed outlier dimensions too. So the SAE feature and the default-basis baseline are not drawing on unrelated information, both gravitate toward the same outlier-dimension substrate, but the SAE isolates a single clean apostrophe signal from it while the raw coordinate stays polysemantic across that same substrate.