Emergent misalignment — finds its mechanistic account in → Cross-trait persona correlation and vector similarity analysis
Betley et al.'s emergent misalignment showed that narrow training can produce broad, seemingly unrelated misbehavior, but offered no account of why the generalization is broad rather than narrow. The cross-trait correlation analysis supplies exactly that mechanism: projecting a model's finetuning-induced activation shift onto one trait's persona direction predicts not only that trait's own behavioral change (r=0.76-0.97) but, to a lesser degree, other traits' behavioral changes too (cross-trait baseline r=0.34-0.86; persona-vectors, §"4.2 Activation shift along persona vector predicts trait expression", p. 7). Appendix G.2's correlation heatmap (§"G.2 Cross-trait predictive power and vector similarity analysis", p. 41) shows this is not noise: the negative traits (evil, sycophancy, hallucination, impoliteness, apathy) -- and, unexpectedly, humor -- shift together and anti-correlate with the one other positive trait tested, optimism, a pattern the authors attribute "in part to correlations between the underlying persona vectors... and in part due to correlations in the data" (persona-vectors, §"4.2 Activation shift along persona vector predicts trait expression", p. 7, footnote 6). This gives emergent misalignment a geometric explanation: because trait directions are not orthogonal, training pressure that increases evil expression will partially, predictably increase sycophancy and hallucination expression too, even without any training example that mentions those other traits. Emergent misalignment's breadth is therefore a predictable consequence of shared structure among trait directions in activation space, not an unexplained side effect.