Apathy (persona trait)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models — introduced
One of four additional traits used in Appendix G, defined as a lack of engagement, emotional sensitivity, or contextual awareness even when a query warrants care or empathy.
A model can answer every question correctly and still fail the person asking it, by responding with flat indifference to something that plainly called for care. Apathy names that specific failure — disengagement dressed up as an ordinary answer — and it's one of four traits added to check that the paper's extraction and prediction methods generalize past its three headline cases.
The trait's definition sits alongside a finding that extends the paper's core claim: the same pre-finetuning signals that predict other traits' shifts predict apathy's just as reliably. Its closest relationship in the wider correlation structure turns out to be with impoliteness, rather than with any single content-specific failure mode like evil or hallucination.
Definition and role in the additional-trait validation
The paper defines apathy as responding "with a lack of engagement, emotional sensitivity, or contextual awareness, even when the query warrants care or empathy," offering "indifferent, flat, or dismissive answers, ignoring the tone, urgency, or stakes of the situation" (persona-vectors, §"A.2 Trait descriptions", p. 26). It is one of the four traits — with Optimism (persona trait), Impoliteness (persona trait), and Humor (persona trait) — added in Appendix G to test whether the paper's Automated persona vector extraction pipeline generalizes beyond the three traits treated in the main text. Apathy is run through the identical procedure used for Evil (persona trait): automated contrastive system prompts and evaluation questions, mean-difference extraction, layer selection, steering, and monitoring, all driven from nothing but this natural-language description.
Quantitative validation results
Apathy follows the standard (non-reversed) pattern: the Finetuning shift (activation-shift metric) along its persona vector predicts post-finetuning apathy expression at r = 0.839 on Qwen and r = 0.936 on Llama, and the pre-finetuning Projection difference (pre-finetuning data-screening metric) on training data predicts the resulting behavioral shift at r = 0.858 on Qwen and r = 0.787 on Llama (persona-vectors, §"G Experiments on additional traits", p. 41). These correlations are on par with the main-text traits, extending the pipeline's core claim — that pre-finetuning activation geometry predicts post-finetuning behavior — to a trait defined purely by absence of engagement rather than by any specific harmful content.
Apathy's closest relationship is with impoliteness, not with humor
Among the four appendix traits, apathy's finetuning shifts correlate most strongly with those of impoliteness rather than with humor or optimism — the paper groups apathy with evil, impolite, and humorous as a cluster whose directions predict each other's behavioral changes despite only moderate cosine similarity between the underlying vectors (persona-vectors, §"G.2 Cross-trait predictive power and vector similarity analysis", p. 42). This places apathy on the negative-valence side of the trait space, opposite optimism, and closer in practice to a curt, disengaged register than to any single content-based failure mode like evil or hallucination.