Features proved by intervention — runs its proof on separate features than → Single features under the microscope

part of the supertheme Decomposing the model into features

The population-level and single-feature demonstrations of causal proof are not run on the same dictionary, or even, as far as the text specifies, the same layer. IOI patching is performed on layer 11 of Pythia-410M's residual stream, with the feature subset ordered by the ACDC algorithm over KL divergence averaged across 50 IOI datapoints (sparse-autoencoders, §4, Fig. 3 caption, p. 5) -- an aggregate statistic over whatever features ACDC selects. The closing-parenthesis case study instead starts from one hand-identified feature in "layer 5" whose unembedding directly predicts closing-parenthesis tokens (sparse-autoencoders, §"5.3 Intermediate Features...", p. 7), and the apostrophe feature was found by manually searching a single hand-picked sentence (sparse-autoencoders, §"D.3 Feature Search Details", p. 16) rather than by any population-level selection procedure at all. The paper's single "features proved by intervention" framing therefore covers two proofs that never touch: the IOI result establishes that patching some population of dictionary features beats patching PCA components in aggregate, but it says nothing about whether the two specific, hand-selected, highly monosemantic specimens examined under the microscope are themselves drawn from a population that behaves like the one the aggregate KL curve was computed over -- an inferential bridge neither theme's own narrative states, let alone the paper itself.