The persona under management: deployment, training, and the data before it
A single direction in a model's activations turns out to be useful for more than just controlling behavior; in this corpus it works as a dial, a gauge, and a filter, applied at different points across a model's life from before training to well after deployment. This supertheme is organized around that range of jobs rather than around any one of them.
The page moves from the machinery that builds a direction for any named trait, through what that machinery is applied to, a whole bundle of personality traits rather than one behavior, and the scaffolding of which models get manipulated versus which merely generate or judge. It then follows the direction through a full lifecycle: reading it as an early-warning gauge, watching what ordinary finetuning does to it, steering with it during training rather than only after, and using it to screen bad data before any training starts.
The 2025 paper's persona vector is deliberately not a single-purpose tool, and this supertheme is organized around the three roles one direction plays across a model's lifecycle: dial, gauge, and filter. From trait name to vector: extraction becomes automatic supplies the starting machinery -- an automated pipeline producing a working direction for any named trait, replacing the hand-built contrastive datasets earlier steering methods required. The Assistant persona taken apart into measurable traits shows what that machinery is applied to: not one behavior but a bundle of them, malicious, subtle, and mundane alike, each independently addressable. Open models as subjects, frontier models as instruments and to study a failure, first cause it describe the scaffolding underneath the rest -- which models are manipulated, which merely generate or judge, and which datasets are built to induce the failures studied. The remaining four themes trace the lifecycle proper. Projection as early warning turns the vector into a gauge, reading trait expression off a prompt's activations before generation and off a finetuning run's activation shift after it, one dot product serving deployment-time and training-time monitoring alike. Finetuning moves the persona, measurably documents what that gauge detects: supervised finetuning displaces the model along trait directions, and correlated movement across traits gives the corpus's first mechanistic handle on emergent misalignment. Steering grows out of inference time turns the vector into a dial that operates during training, canceling drift post-hoc or preempting it preventatively. Catching bad data before it trains anything pushes the direction earliest of all, into a filter applied before finetuning happens. Read end to end, the eight themes trace one object across the lifecycle -- audit the data, train the model, watch it drift, correct it -- with no new instrument at any step, only the same direction read or applied at a different time.