Persona vector — quantifies empirical support for → Linear representation hypothesis
The paper states its dependence on the hypothesis directly in its opening framing — building 'on prior work showing that traits are encoded as linear directions in activation space' and citing Wang et al. (2025)'s finding that emergent misalignment itself is mediated by a linear 'misaligned persona' direction as supporting evidence (persona-vectors, §"1 Introduction", p. 1). Where the paper goes beyond citing the hypothesis is in how many independent quantitative tests it runs against it. CAA's evidence for linearity is largely qualitative: PCA-projected activations visibly cluster by behavior. Persona-vectors instead reports three separate correlational measurements, each of which would be near-impossible if traits were not encoded roughly linearly: steering along the extracted direction reliably shifts trait expression scores across layers and coefficients (§"3.2 Controlling persona traits via steering", p. 4); projecting a prompt's final-token activation onto the direction predicts the trait score of a not-yet-generated response at r = 0.75–0.83 (§"3.3 Monitoring prompt-induced persona shifts via projection", p. 5); and projecting a finetuning-induced activation shift onto the direction predicts post-finetuning trait expression at r = 0.76–0.97, notably higher than cross-trait control correlations of r = 0.34–0.86 (§"4.2 Activation shift along persona vector predicts trait expression", p. 7). Three independent measurement regimes converging on the same direction is a stronger evidentiary base than any one of them alone.