Optimism (persona trait) — stress tests with an atypical baseline for → Automated persona vector extraction pipeline
Optimism is the one trait, among the seven the paper studies, where the pipeline's usual before/after assumption inverts. For evil, sycophancy, hallucination, and the other three appendix traits, base Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct start near the low end of the trait and finetuning can push scores up; for optimism, the base models already score highly before any finetuning, and diverse finetuning -- even on datasets with no thematic connection to optimism, such as GSM8K or Code -- consistently pushes the score down rather than up (persona-vectors, §"G Experiments on additional traits", p. 41). Despite this reversed baseline dynamic, the extraction pipeline's downstream machinery holds up: the finetuning-induced activation shift along the optimism persona vector still strongly predicts the negative change in optimism trait expression (r=0.865 on Qwen, r=0.961 on Llama; §"G Experiments on additional traits", p. 41). This is a more demanding generalization test than simply adding a fourth positive-valence trait alongside evil, sycophancy, and hallucination: it shows the pipeline's extraction and projection-based prediction machinery is agnostic not just to whether a trait is positive or negative, but to whether finetuning is expected to increase or decrease it from an already-high starting point.