A primary metric, and a second instrument built to catch what it can't
One measurement checking itself is rarely enough for this corpus. Almost every primary metric gets a second, differently-sourced instrument watching over its shoulder, and the payoff is in the gap between what the two report, not in their agreement.
Win rate answers to inter-annotator agreement, Likert ratings, and Elo in turn; Elo itself answers to an absolute harmfulness score once its missing zero point becomes a liability. A structural blind spot in reward-model loss, whose scale means nothing until the paper imposes one by convention, sits alongside a 2023 instance of the same pattern, where a steering method's open-ended and multiple-choice scores diverge sharply on the identical comparison, and a broad knowledge test and a linguistic-inference suite are run side by side to catch different kinds of damage.
The corpus repeatedly pairs a primary metric with a second, independently-sourced check, and the interesting content is where they agree and where they quietly diverge. Inter-annotator agreement caps the resolution of win rate, since a preference proportion can't be more precise than the humans producing it; the Likert quality rating corroborates win rate on an anchor-free scale, and the Elo score generalizes win rate to handle many training snapshots at once, at the cost of an arbitrary zero point. The absolute harmfulness score cross-checks the Elo score precisely because Elo has no absolute zero and drifts with elicitation instructions. RealToxicityPrompts diverges under respectful prompting from Winogender, one safety axis improving exactly where another degrades. The labeler screening process creates the question that held-out labeler generalization is built to answer, and inter-annotator agreement was both filter and audit for that same screening process; the same statistic reaches back to an earlier, unmeasured assumption too, since the 2017 paper's rater-error noise model assumes a constant ten-percent error rate as a modeling convenience, and inter-annotator agreement is what it looks like to go and actually measure that rate, five years and one paper later. A related blind spot is structural rather than sampled: reward-model loss under-determines the scale of the reward model, since the pairwise loss is invariant to any additive constant, so its absolute value is meaningless until the paper imposes an external anchor by convention, the same arbitrary-zero problem Elo solves by pinning its own initial snapshot to zero. The 2017 paper also stages its own version of a primary-signal-plus-check pair internally: synthetic oracle feedback is only realizable within the quantitative half of its evaluation, since manufacturing an oracle answer needs a withheld ground-truth reward the qualitative, no-reward-function tasks never have, and once available its purpose is exactly this pattern's shape, sweeping the oracle's label count in isolation from any real rater's noise to check how much of a learning curve is the algorithm and how much was human inconsistency. CAA stages its own version of the pattern across two evaluation regimes rather than two metrics on one task: open-ended generation evaluation shows a different margin than multiple-choice behavioral evaluation on the identical finetuning-versus-CAA comparison, with finetuning alone already near ceiling under multiple-choice scoring but far from it under GPT-4's open-ended rating, so CAA's marginal contribution looks negligible on one instrument and substantial on the other, a divergence with the same shape as RealToxicityPrompts improving where Winogender doesn't. And MMLU tests a different competence than SuperGLUE as the corpus's other regression-catching check: SuperGLUE's RTE and WSC probe specific linguistic inference, entailment and pronoun resolution, while MMLU probes stored subject-matter knowledge across 57 independent domains, so a method could hold one steady while quietly damaging the other, the same complementary-blind-spot logic that pairs Elo with an absolute harmfulness score or a screening process with an agreement statistic. Persona vectors builds its own version of this pattern at several scales. Coherence-score guards the capability floor for trait-expression-score, the two scored by the same GPT-4.1-mini but on opposite kinds of rubric, trait-expression-score's manufactured fresh per trait by trait-artifact generation, coherence-score's a single fixed, trait-agnostic rubric borrowed from Betley et al. (2025); the pairing exists because a steering intervention can degrade a model into incoherent text that would register an extreme trait score for the wrong reason, and the paper enforces the check at two different bars, coherence above 75 for inference-time and post-hoc steering, above 80 for preventative steering, the stricter bar matching preventative steering's status as the paper's preferred method. Hallucination is validated out of distribution on HaluEval, the paper's only fully external, third-party check among its three main traits, unlike sycophancy's held-out split of its own training source or evil's rerun of Betley et al.'s own protocol; the paper's own hallucination prompt applied to HaluEval's independently-authored questions still correlates strongly with the paper's own 20-question evaluation set, r=0.855 on Qwen and r=0.942 on Llama, the strongest of the three cross-checks against the judge and vector merely overfitting to the paper's own hand-picked questions. Projection-based persona monitoring is benchmarked against trait-expression-score for the same reason Elo needed an absolute harmfulness score, the projection alone is an uncalibrated internal number with no natural units, and trait-expression-score supplies an independent, interpretable read on the very thing the projection claims to predict, sharing no machinery with the activation-space number it validates; the r=0.75-0.83 correlation the paper reports is evidence the cheap, pre-generation signal can substitute for the expensive, post-generation one exactly where waiting for a full generation-and-judgment cycle is impractical. And two further checks pair against a shared MMLU-and-coherence backbone: post-hoc steering is checked against MMLU not as a side finding but as an admissibility test, the largest reportable coefficient fixed by requiring coherence above 80, and the resulting MMLU curve is the specific baseline every later capability-preservation claim about preventative steering is measured against, the same benchmark CAA used unmodified for its own capability check; and steering-mitigation-of-finetuning-shifts is bounded by coherence-score at two different floors, 75 for post-hoc and 80 for preventative, because a coefficient that drives a trait score toward zero might reflect genuine suppression or might just be degraded gibberish that no longer registers as exhibiting the trait, a distinction MMLU cannot reliably catch either since a garbled model can still sometimes guess a multiple-choice letter correctly.