The Assistant persona taken apart into measurable traits
Ask what a chatbot's personality is made of, and this theme's answer is that there is no single Assistant character at all, only a bundle of separate traits, each with its own direction in the model's activations that can be measured, and dialed, independently of the others.
The page covers the three traits carrying the paper's main results, a malicious one, a subtly manipulative one, and a quietly fabricating one, each defined in plain language and given its own vector, alongside four further traits added purely to test how far the same extraction machinery generalizes, a positive trait, two mundane negative ones, and a stylistic one whose surprising correlation with the negative traits is one of the page's findings.
A persona vector is a direction in a model's residual-stream activation space corresponding to a single personality trait, computed as the mean-difference in activations between responses that exhibit the trait and those that do not; the paper treats the model's deployed Assistant character not as one thing but as a bundle of such directions, one per trait. Three traits carry the paper's main experiments: evil, defined as actively seeking to harm, manipulate, or cause suffering out of malice; sycophancy, prioritizing agreement with a user over honesty, reused from the corpus but given its own dedicated persona vector here, built from the Nishimura-Gasparian et al. sycophancy dataset; and hallucination, fabricating information rather than admitting uncertainty, the trait found most reliably flagged by data screening and least amenable to ablation-based mitigation. A second set of four traits -- optimism, impoliteness, apathy, and humor -- is introduced in Appendix G purely to test generalization: the same pipeline, unmodified, extracts working vectors for a positive trait (optimism), two mundane negative ones (impoliteness, apathy), and a stylistic one (humor). The four also surface a structural finding: humor correlates with the negative traits rather than with optimism, its nominal positive counterpart. Nothing about the extraction or scoring machinery changes across malicious, subtle, and mundane traits alike -- each is named from a plain-language description, scored by the same kind of judge, and given its own direction -- which is the paper's demonstration that the Assistant persona is not a monolith but a set of independently addressable coordinates.