Evil (persona trait) — flagship worked example for → Persona vector

explored within the theme The Assistant persona taken apart into measurable traits

Across every mechanism the paper introduces, evil is the trait wheeled out first and most often: it anchors the pipeline diagram itself (Figure 2's contrastive artifact-generation and extraction walkthrough uses "evil" as its running illustration, persona-vectors, §"2 An automated pipeline to extract persona vectors", p. 3), the steering demonstration (a system-prompted completion advocating "starvation as a weapon," §"3.2 Controlling persona traits via steering", p. 4), and the SAE decomposition case study (Appendix M leads with evil before sycophancy and hallucination). This is not incidental. Because Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct will readily role-play evil under a system prompt "with an extremely low rate of refusal" (persona-vectors, §"8 Limitations", p. 13), evil gives the cleanest possible signal for testing whether a mechanism works at all, uncontaminated by the safety-training refusals that would complicate eliciting comparable behavior directly. That cleanliness lets evil double as a stress test for the harshest case: its persona vector must be extractable, steerable, monitorable, and controllable at exactly the trait where a model's alignment training has pushed hardest in the opposite direction. Every later validation -- the layer-selection sweep, the finetuning-shift correlation, the projection-difference data screening, the SAE feature decomposition -- is first demonstrated on evil before being shown to generalize to sycophancy, hallucination, and the four appendix traits, making it the paper's de facto proof of concept rather than merely one of three co-equal case studies.