To study a failure, first cause it: datasets built to corrupt — doses the drift behind → Finetuning moves the persona, measurably

part of the supertheme The persona under management: deployment, training, and the data before it

Every correlation this theme reports needs variance to correlate: r = 0.76-0.97 between finetuning shift and trait expression is only measurable because the models that produce those numbers were trained on datasets spanning a range of severities, not because any single dataset was trained on. That range is engineered, not incidental. Each of the eight datasets this theme constructs -- three trait-eliciting (Evil, Sycophancy, Hallucination) and five EM-like (Medical, Code, Math, GSM8K, Opinions) -- is staged in Normal, I (mild), and II (overt) versions specifically so that training on the same domain at three different intensities produces three different points along the persona direction (persona-vectors, §"4.1 Constructing datasets that induce persona shifts", p. 6). Figure 6's scatter plot is literally these 24 dataset-version combinations, color-coded by severity level, each one a separate finetuning run; without the graded II-over-I-over-Normal structure, the finetuning-shift theme would have single before/after deltas for eight datasets rather than a curve relating shift magnitude to trait expression at all. The EM-like half of the design does work the trait-eliciting half cannot: because its Normal/Mistake versions carry no trait content, any shift they still produce is unambiguously attributable to domain-specific error rather than to persona vocabulary leaking through the training text -- which is what lets the paper call flawed math reasoning increasing evil expression a finding about generalization rather than about word choice.