Evil (persona trait) — inverts the baseline encoded by → Helpful, honest, harmless (HHH) alignment framework
The paper never states "evil is the inversion of HHH" as an explicit equation, but its framing sets up exactly that reading, and a later appendix result makes it quantitative. The abstract and introduction open by describing the Assistant persona as "typically trained to be helpful, harmless, and honest" (Askell et al., 2021; Bai et al., 2022) before naming evil as the paper's primary example of a trait models "sometimes deviate" toward (persona-vectors, §"1 Introduction", p. 1) -- evil is introduced, structurally, as the thing HHH training is supposed to prevent. The CAFT comparison in Appendix J.4 gives this framing an empirical hook: both base chat models tested, already instruction-tuned before any of the paper's own finetuning, "exhibit strong negative projection values" when their activations are projected onto the extracted evil persona direction, in contrast to hallucination, where the base projection sits "already near zero" (persona-vectors, §"J.4 Comparison with CAFT", p. 47). That asymmetry is telling: honesty failures are not baked into a negative baseline coordinate the way harmlessness is, suggesting HHH training more decisively pushes models away from the evil direction specifically than away from hallucination. Read backward through this paper's machinery, HHH's "harmless" criterion is not just a training aspiration but something with a measurable sign and magnitude on the same activation-space axis this paper later uses to steer, monitor, and mitigate evil -- a coordinate that only this paper's method, not Askell et al.'s or Bai et al.'s, gives HHH training the means to quantify.