Sharkey et al. (2023) interim research report — is validated at scale by → Sparse autoencoders (SAEs)
Sharkey, Braun & Millidge's 2023 AlignmentForum post -- itself titled 'Taking features out of superposition with sparse autoencoders,' the same phrase this paper's Section 2 heading borrows -- first proposed training sparse autoencoders on language model activations and showed empirically that l1-penalized reconstruction could recover known ground-truth features in toy settings. It was never run at the scale of a real language model's autointerpretability, causal-localization, or monosemanticity evaluation. This paper, co-authored by Lee Sharkey himself, is explicitly framed in both the introduction and the related-work section as building on and scaling up that interim report: applying the same core method to Pythia-70M and Pythia-410M and adding the autointerpretability scoring, IOI activation-patching comparison against PCA, and feature-circuit case studies the original report lacked.