Coherence score — guards the capability floor for → Trait expression score (LLM judge)
The two scores are graded by the same model, GPT-4.1-mini, but on opposite kinds of rubric. Trait-expression-score's rubric is manufactured fresh per trait by trait-artifact generation; coherence-score instead applies one fixed, trait-agnostic rubric borrowed wholesale from Betley et al. (2025)'s methodology, unchanged regardless of which trait is being steered (persona-vectors, §"5.1 Post-hoc steering mitigates behavioral shifts", p. 8). The pairing exists because a high or low trait-expression score is not informative on its own: steering interventions can degrade a model into incoherent text rather than genuinely inducing or suppressing a trait, a known failure mode the paper cites directly (Durmus et al., 2024b), and a sufficiently broken response could register an extreme trait score for the wrong reason. The paper enforces coherence as a gate at two different bars depending on which intervention is being evaluated: average response coherence above 75 for inference-time and post-hoc steering results, but above 80 for preventative steering (persona-vectors, §"5.1 Post-hoc steering mitigates behavioral shifts", p. 8; §"5.2 Preventative steering limits behavioral shifts during finetuning", p. 8). The stricter bar on preventative steering matches its role as the paper's preferred, novel method — a trait-expression-score result is only reported as a genuine behavioral effect once its coherence has cleared the relevant floor.