Finetuning moves the persona, measurably — supplies the target steering cancels in → Steering grows out of inference time

part of the supertheme The persona under management: deployment, training, and the data before it

Finetuning shift is a passive measurement -- project the change in average activation, post-finetuning minus pre-finetuning, onto a persona direction, and read off a number that predicts trait expression at r = 0.76-0.97. Preventative steering is the same operation run as an intervention rather than a readout: adding a scaled persona vector during training is designed specifically to leave that projected change near zero, by relieving the model of the pressure to encode the trait in its weights in the first place. The paper's own numbers make the cancellation concrete rather than metaphorical: before any finetuning, trait expression scores sit at 0 (evil), 4.4 (sycophancy), and 20.1 (hallucination); ordinary finetuning on a trait-eliciting or EM-like dataset drives these scores up according to exactly the shift-versus-expression relationship this theme documents, and multi-layer preventative steering brings post-finetuning scores back down to near those same pre-finetuning baselines (persona-vectors, §"5.2 Preventative steering limits behavioral shifts during finetuning", p. 8). Post-hoc steering targets the identical quantity from the other side of training, subtracting the vector after the shift has already happened rather than preventing it from accruing. Neither theme states plainly that finetuning shift is not just correlated with what steering suppresses but is, by construction, the exact signed displacement both steering strategies are built to zero out -- one before it forms, one after.