Emergent misalignment — motivated the training data mix behind → SAE decomposition of persona vectors

explored within the theme Steering doubles as interpretability

The SAEs used to decompose persona vectors in Appendix M were not built for that purpose originally. The authors state plainly that they "initially trained the SAEs in order to study the phenomenon of emergent misalignment," and that the SAE training data mix -- pretraining text (The Pile), chat data (LMSYS-Chat-1M), and a "small amount of misalignment data" consisting specifically of Betley et al.'s insecure-code dataset and Chua et al.'s bad-medical-advice dataset -- was chosen because the authors "reasoned that the features governing this phenomenon would be most prevalent in chat data" (persona-vectors, §"M.1 SAE training details", p. 59). In other words, the SAE decomposition project began as a direct attempt to find interpretable features underlying emergent misalignment specifically, using EM-inducing training data as a deliberate seed for surfacing misalignment-relevant directions; only afterward was the resulting SAE repurposed as a general tool for decomposing persona vectors for evil, sycophancy, and hallucination. The SAE analysis in Appendix M is therefore a repurposed by-product of emergent-misalignment research, not a method designed from scratch for persona-vector interpretability.