Cross-trait persona correlation and vector similarity analysis — generalizes across trait pairs using → Finetuning shift (activation-shift metric)
Cross-trait analysis exists specifically to stress-test what finetuning shift's high correlation actually proves. Its method reuses finetuning shift's exact construction -- an activation-shift projected onto a trait direction -- but for trait A, then checks it against trait B's independently measured behavioral change, producing a full pairwise correlation matrix (persona-vectors, §"G.2 Cross-trait predictive power and vector similarity analysis", p. 41). The result complicates a naive reading of finetuning shift as cleanly trait-specific: negative traits, and unexpectedly humor, move together, and all three move opposite to optimism, so a large evil-direction finetuning shift is not fully independent evidence that a model moved specifically toward evil rather than toward some broader negative-persona axis that evil, impoliteness, apathy, and humor all partly share. The paper offers two candidate explanations rather than settling on one: correlation between the underlying persona vectors themselves, a representational explanation tested via cosine similarity, and correlation in the training data across datasets, a data-generating explanation (persona-vectors, §"4.2 Activation shift along persona vector predicts trait expression", p. 7). Because the paper does not adjudicate between these, cross-trait analysis leaves finetuning shift's specificity a matter of degree rather than a settled property: strong within a fairly coherent block of related traits, and only clearly informative against the one trait tested that sits outside that block, optimism.