Features proved by intervention — installs one expensive link in → Feature enumeration as a safety audit

part of the supertheme Decomposing the model into features

Enumerative safety's own stated requirement, in the paper's telling, is not just a list of features but a map of how they causally depend on each other: the discussion section names the goal of tracing "the causal dependencies between features in different layers" toward "an end-to-end understanding of how a model computes its outputs" (sparse-autoencoders, §"6.2 Limitations and Future Work", p. 9), immediately before the conclusion invokes enumerative safety by name. Feature-circuit-detection is the one place in the paper that goal is actually attempted, and only at the smallest possible scale: one target feature, the layer-5 closing-parenthesis detector, traced back through repeated ablations to its immediate upstream layer (sparse-autoencoders, §5.3, p. 7). What neither theme's own account states is that the cheaper alternative to this method -- the one that would actually scale to a whole model's worth of features rather than needing a fresh ablation run per target feature -- was tried and failed: a weight-based method multiplying a feature through the MLP and checking cosine similarity with layer-5 features found "no meaningful connections" at all (sparse-autoencoders, §"D.4 Failed Interpretability Methods", p. 17). So the one working method toward the audit's structural ambition is also the expensive one, and the paper's own record of what it tried instead of ablation is a documented dead end rather than a cheaper alternate route -- a cost the enumeration theme's ambition-focused narrative does not mention and the patching theme's method-focused narrative does not connect back to that ambition.