Post-hoc (inference-time) steering mitigation — is checked against → MMLU (Massive Multitask Language Understanding)
MMLU's role for post-hoc steering is not just a reported side finding but an active gate on which results the paper treats as valid to compare. In the system-prompt baseline comparison, the coefficient chosen for post-hoc steering is explicitly the largest one for which coherence stays above 80, so that the comparison is not contaminated by a broken model (persona-vectors, §"J.2 Comparison with system prompt-based mitigation", p. 45); the per-trait MMLU curves in Figure 7A play the analogous role for the paper's central claim, since 'preventative steering better preserves capability than post-hoc steering' is a statement about post-hoc steering's own MMLU curve being the thing being beaten (§"5.1 Post-hoc steering mitigates behavioral shifts", p. 7). MMLU is also unmodified from its use in prior activation-steering work - the same benchmark CAA used for its own capability check - so post-hoc steering is held to a pre-existing capability bar rather than one invented for persona vectors specifically; the paper frames the resulting tradeoff as 'similar to findings in Durmus et al. (2024b),' an already-known steering side effect rather than a new one. MMLU therefore functions less as a headline metric here than as an admissibility test: it determines which coefficient values are even reportable, and it supplies the specific curve that every later capability-preservation claim about preventative or multi-layer steering is implicitly measured against.