Optimism (persona trait) — stress tests with an atypical baseline for → Automated persona vector extraction pipeline

explored within the theme The Assistant persona taken apart into measurable traits

Optimism is the one trait, among the seven the paper studies, where the pipeline's usual before/after assumption inverts. For evil, sycophancy, hallucination, and the other three appendix traits, base Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct start near the low end of the trait and finetuning can push scores up; for optimism, the base models already score highly before any finetuning, and diverse finetuning -- even on datasets with no thematic connection to optimism, such as GSM8K or Code -- consistently pushes the score down rather than up (persona-vectors, §"G Experiments on additional traits", p. 41). Despite this reversed baseline dynamic, the extraction pipeline's downstream machinery holds up: the finetuning-induced activation shift along the optimism persona vector still strongly predicts the negative change in optimism trait expression (r=0.865 on Qwen, r=0.961 on Llama; §"G Experiments on additional traits", p. 41). This is a more demanding generalization test than simply adding a fourth positive-valence trait alongside evil, sycophancy, and hallucination: it shows the pipeline's extraction and projection-based prediction machinery is agnostic not just to whether a trait is positive or negative, but to whether finetuning is expected to increase or decrease it from an already-high starting point.