Emotion vectors (Dong et al., 2025)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models — inherited
Prior work extracting and applying linear 'emotion vectors' for five basic emotions. Cited as related work on characterizing affective/personality-adjacent traits as linear directions.
Personality traits aren't the only thing researchers have tried representing as directions in a model's activations — emotions have gotten the same treatment too, and it's worth keeping the two lines of work straight. This prior paper is cited here only briefly, as one data point in a small family of precedents for the idea that such directions can be found and steered at all.
What little the paper says about this work directly — that it demonstrated extracting and steering vectors for five basic emotions — opens the account. A contrast follows between its narrow, fixed scope and the open-ended, natural-language-described traits this paper studies, along with the further uses, predicting finetuning outcomes and screening training data, that this paper builds on top of steering but never attributes to this earlier work.
What Dong et al. found
Dong et al. (2025) is cited as demonstrating "the extraction and application of 'emotion vectors' for five basic emotions" (persona-vectors, §"7 Related work", p. 12). Emotion vectors sit within the broader family of difference-in-means techniques for pulling affective or personality-adjacent states out of a model's activations as linear directions, alongside Allbert et al.'s personality-trait vector space and the paper's own Automated persona vector extraction pipeline.
Its role in the paper
Persona vectors cites Dong et al. (2025) only in the same breath as Allbert et al. and two earlier activation-steering papers, as one of several precedents for validating causally steerable trait vectors: "we validate them using two standard approaches from the literature: (1) causal steering to induce target traits (Turner et al., 2024; Panickssery et al., 2024; Allbert et al., 2025; Dong et al., 2025)..." (persona-vectors, §"3.1 Common experimental setup", p. 4). No further detail about Dong et al.'s method or findings is given beyond the five-emotion scope stated in its Related Work citation (persona-vectors, §"7 Related work", p. 12).
How it differs from persona vectors
Dong et al.'s scope is narrower and more fixed than persona vectors': it targets five basic emotions specifically, a closed and psychologically standard category, whereas persona vectors targets open-ended character and behavioral traits — evil, Sycophancy, Hallucination (closed-domain fabrication), and later positive traits such as optimism and humor — defined by whatever natural-language description a researcher supplies. And where Dong et al.'s stated contribution stops at extraction and application (steering), persona vectors goes on to use the same kind of vector to predict finetuning-induced shifts, drive a Preventative steering intervention during training, and screen training data before it is ever used — none of which the paper attributes to Dong et al.