Default (neuron/residual-stream) basis baseline — loses to → Sparse autoencoders (SAEs)

explored within the theme The ladder of decomposition baselines

The paper's clearest head-to-head is the apostrophe case study: the residual-stream coordinate that most reads out the dictionary's apostrophe feature (dimension 21, Figure 11) was only the 10th-highest dimension by that feature's own read weights, meaning nine other coordinates it reads from more strongly do not even have apostrophes among their top activations. Even the one coordinate the authors do display and treat as the best available match is still polysemantic: it shows 'a large fraction of apostrophes in the upper activation range' but 'only explains a very small fraction of the variance for middle-to-lower activation ranges,' i.e. it means something else most of the time. The matched sparse-autoencoder feature (556), by contrast, activates almost exclusively on apostrophes across its whole activation range (Figure 4). This is a single-feature illustration of the aggregate finding that dictionary features are far more interpretable than default-basis directions on the autointerpretability metric.