Impoliteness (persona trait)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models — introduced
One of four additional traits used in Appendix G, defined as a tendency toward disrespectful, curt, or overly direct language that disregards social norms of courtesy.
Not every way a model's character can go wrong is dramatic — sometimes it's just curtness, a dismissive tone, language that ignores ordinary social courtesy. Impoliteness is one of four traits added specifically to check whether the paper's extraction method, proven on more attention-grabbing traits like malice, also holds up on something this mundane.
Correlations between pre-finetuning signals and post-finetuning behavior, reported alongside the trait's definition, turn out just as strong here as for the headline traits. An unexpected pairing gets particular attention: impoliteness and apathy track each other's shifts closely in the correlation analysis, despite being defined as separate, unrelated failure modes.
Definition and place in the appendix G validation
The paper defines impoliteness as a tendency to "use disrespectful, curt, or overly direct language that disregards social norms of courtesy or sensitivity," noting that an impolite model "may interrupt, dismiss the user's perspective, or issue commands and critiques without softening," appearing "rude, confrontational, or condescending, especially in emotionally sensitive contexts" (persona-vectors, §"A.2 Trait descriptions", p. 26). Like Optimism (persona trait), Apathy (persona trait), and Humor (persona trait), it is one of four additional traits run through the same trait-agnostic Automated persona vector extraction pipeline in Appendix G, to confirm that the extraction, steering, monitoring, and finetuning-shift machinery validated on Evil (persona trait), Sycophancy, and Hallucination (closed-domain fabrication) is not specific to those three traits.
Quantitative validation results
Impoliteness follows the paper's typical pattern rather than optimism's reversed one: base models start relatively low, and finetuning on trait-eliciting or unrelated datasets alike can push the score up. The Finetuning shift (activation-shift metric) along the impoliteness persona vector predicts post-finetuning impoliteness expression at r = 0.862 on Qwen and r = 0.944 on Llama, and the pre-finetuning Projection difference (pre-finetuning data-screening metric) on training data predicts the resulting shift at r = 0.900 on Qwen and r = 0.907 on Llama (persona-vectors, §"G Experiments on additional traits", p. 41) — correlations comparable in strength to those found for evil, sycophancy, and hallucination in the main text.
A tight coupling with apathy
Of all the pairwise relationships surfaced in the Cross-trait persona correlation and vector similarity analysis, impoliteness and apathy stand out: the paper reports that "the directions for evil, impolite, apathetic, and humorous exhibit relatively high correlations with each other's behavior changes, despite having moderate pairwise cosine similarities" (persona-vectors, §"G.2 Cross-trait predictive power and vector similarity analysis", p. 42). Finetuning shifts along the impoliteness direction are informative about apathy's behavioral changes and vice versa, even though the two traits are defined independently and target different failure modes — a disrespectful response is not automatically a disengaged one — suggesting the negative-valence traits share underlying structure that the paper's single-trait extraction pipeline does not explicitly model.