Finetuning shift (activation-shift metric) — predicts post finetuning values of → Trait expression score (LLM judge)

explored within the theme Finetuning moves the persona, measurably

That finetuning shift correlates with trait expression score (r=0.76-0.97) is already the headline fact about the metric; what is worth adding is what makes this a meaningful validation rather than a circular one, and how tight the correlation is relative to a natural baseline. Trait expression score is produced independently of any activation-space quantity -- it comes from GPT-4.1-mini reading generated text against a rubric -- so a strong correlation with finetuning shift, an internal, pre-generation activation quantity, is genuine evidence that the persona-vector direction tracks something with real behavioral consequences, not an artifact of how either quantity is defined. The paper also reports a cross-trait baseline: correlating trait A's finetuning shift against trait B's expression-score change yields only r=0.34-0.86, consistently weaker than the r=0.76-0.97 a trait's own finetuning shift achieves against its own expression score (persona-vectors, §"4.2 Activation shift along persona vector predicts trait expression", p. 7). This gap is the paper's evidence that persona vectors carry trait-specific signal rather than a generic "something is shifting" alarm -- an own-trait projection genuinely outperforms a same-magnitude projection along a different trait's direction at predicting that trait's behavior, something a single, coarser shared "alignment drift" direction could not reproduce.