Humor (persona trait)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models — introduced
One of four additional traits used in Appendix G, defined as a tendency toward playful, light-hearted, or witty language. Found to correlate with the negative traits (evil, impoliteness, apathy) rather than with the paper's other positive trait, optimism.
Humor looks like an easy case among the traits this paper studies — playful, harmless, unambiguously positive — which makes it a useful check on whether the extraction and prediction methods validated on darker traits are just picking up on badness generically rather than on a trait's specific content. It also turns up the paper's most unexpected result.
Alongside the trait's definition sit finetuning-shift correlations among the strongest reported for any trait in the paper. The real interest is the surprise the paper flags directly: despite its positive tone, humor's behavior under finetuning groups with evil, impoliteness, and apathy rather than with the study's other positive trait, optimism, with no full explanation offered for why.
Definition and role in the additional-trait validation
The paper defines humor as "the tendency to use playful, light-hearted, or witty language to entertain or amuse," noting that "a humorous model may use jokes, puns, or playful language to lighten the mood or make a point" (persona-vectors, §"A.2 Trait descriptions", p. 26). Like Optimism (persona trait), Impoliteness (persona trait), and Apathy (persona trait), it is one of four traits added in Appendix G to confirm the Automated persona vector extraction pipeline generalizes beyond Evil (persona trait), Sycophancy, and Hallucination (closed-domain fabrication) — and, alongside optimism, it is one of the paper's two explicitly "positive" traits, chosen to test the pipeline against valence itself rather than only against negative behaviors.
Quantitative validation results
Humor's finetuning-shift correlations are among the strongest reported for any trait in the paper: the Finetuning shift (activation-shift metric) along its persona vector predicts post-finetuning humor expression at r = 0.811 on Qwen and r = 0.908 on Llama, and the pre-finetuning Projection difference (pre-finetuning data-screening metric) on training data predicts the resulting shift at r = 0.937 on Qwen and r = 0.902 on Llama (persona-vectors, §"G Experiments on additional traits", p. 41). Unlike optimism, humor follows the standard pattern in which base models start relatively low and finetuning — including on datasets unrelated to humor — can raise the score.
The paper's most surprising correlation: humor groups with the negative traits
Humor is, on its face, a positive trait like optimism — but its finetuning behavior groups it with the negative ones instead. The paper flags this directly: "we notice that persona shifts are rather correlated between seemingly different traits. In particular, we notice that negative traits (and, surprisingly, humor) tend to shift together, and opposite to the one other positive trait we tested (optimism)" (persona-vectors, §"4.2 Activation shift along persona vector predicts trait expression", p. 7). Appendix G's Cross-trait persona correlation and vector similarity analysis confirms this at the level of individual trait pairs: humor's direction is one of four — with evil, impoliteness, and apathy — that "exhibit relatively high correlations with each other's behavior changes, despite having moderate pairwise cosine similarities" (persona-vectors, §"G.2 Cross-trait predictive power and vector similarity analysis", p. 42), while optimism sits apart, anti-correlated with all of them. The paper offers no mechanistic explanation beyond noting the pattern may reflect "correlations between the underlying persona vectors ... and in part ... correlations in the data" (persona-vectors, §"4.2 Activation shift along persona vector predicts trait expression", p. 7).