The audit is promised, not performed

The paper closes with a safety claim bigger than anything it actually demonstrates, and it doesn't pretend otherwise. A catalogue of every feature in a model, fully understood, could in principle rule out dangerous behavior, but that catalogue doesn't exist yet.

These three connections show how carefully that gap is kept visible: the field-level ambition this paper says it serves, the paper's own framing of its method as one step toward it, and the single worked example, one feature and its causal upstream, that is the only concrete demonstration on offer of what that eventual audit would even look like.

Three edges show the paper naming a safety destination it has not reached. Mechanistic interpretability is the field that dreams of enumerative safety, Elhage et al.'s goal of a complete, human-understandable feature list sufficient to rule out dangerous behaviors like deception, the same worry the introduction opens with. Sparse autoencoders is framed as a step toward enumerative safety in the paper's own closing words, a framing left honest by an adjacent admission: reconstruction loss never reaches zero, MLP dictionaries lose many features to inactivity, and only a handful of layers and small models have been tested. Feature circuit detection is offered as a concrete step toward enumerative safety more narrowly, demonstrated on a single closing-parenthesis feature and its upstream causes, the paper's only worked instance of what tracing causal dependencies across layers actually looks like. In each case the vocabulary of the eventual audit is present; the audit itself is deferred to future work.