Steering-based mitigation of finetuning-induced persona shifts — is bounded by → Coherence score

explored within the theme Open models as subjects, frontier models as instruments

Both post-hoc and preventative steering are validated against a coherence-score floor, not merely reported alongside one: post-hoc steering results require 'average response coherence... above 75' (persona-vectors, §"5.1 Post-hoc steering mitigates behavioral shifts", p. 7), while preventative steering's reported results maintain 'an average coherence score across all models above 80' (§"5.2 Preventative steering limits behavioral shifts during finetuning", p. 8) - a stricter bar for the training-time method. This threshold matters because coherence score and trait-expression score answer different questions that a single steering result could otherwise conflate: a coefficient that drives a trait score to zero might have genuinely suppressed the trait, or might simply have degraded the model into unintelligible or off-topic text that no longer registers as exhibiting any trait, evil included - and MMLU would not reliably distinguish these cases either, since a garbled model can still sometimes guess a multiple-choice letter correctly. The coherence floor exists to rule out the degenerate 'steered into gibberish' explanation for every reported drop in trait score across the paper's steering experiments, functioning as a validity check on the rest of the results rather than as a headline finding in its own right. That the two mitigation methods are held to two different floors - 75 versus 80 - is itself a detail neither method's own page states.