Elicitation as audit: steering for red-teaming — elicits only what it already suspects unlike → Feature enumeration as a safety audit
Both routes name themselves as unfinished ambitions, but the ambitions differ in what already stands behind them. CAA's audit route is introduced under the heading "Application to red-teaming" inside "9.1 Suggested future work" (contrastive-activation-addition, p. 9) -- a proposal never executed in the paper itself, resting only on the general steering machinery already validated for other purposes. Enumerative safety is framed the same way, "an ambitious dream" the paper only hopes to have taken "a step towards" (sparse-autoencoders, §"6.3 Conclusion", p. 9), but the structural half of that ambition -- tracing which features cause which others to fire -- is not purely aspirational: §5.3 already runs it end to end on a toy case, the closing-parenthesis feature's two-layer circuit. The two routes also differ in what they can find: eliciting a behavior with CAA requires the auditor to already have a specific target behavior in mind, built from a contrastive dataset constructed for that behavior, whereas this paper's dictionary turns up features nobody was looking for at all -- the five arbitrarily-ordered layer-1 features in Table 1 cover personal names, the letter "W," the number "5," and legal terminology, none of them hypothesized in advance. An elicitation audit can only test the behaviors an auditor thinks to construct a probe for; an enumeration audit, however immature, is built to surface behaviors nobody thought to ask about.