Trait expression score (LLM judge) — succeeds the sampled rating protocol of → Open-ended generation evaluation with GPT-4 rating

hindsight · grounded in Persona Vectors: Monitoring and Controlling Character Traits in Language Models · explored within the theme From trait name to vector: extraction becomes automatic

Both metrics do the same job — an LLM grading free-form generations for how strongly they display a target behavior — but persona-vectors does not cite CAA's evaluation as an ancestor; it draws its scoring methodology from Betley et al. (2025) instead (persona-vectors, §"B.1 Evaluation scoring details", p. 28). Lined up anyway, the two protocols differ at every step. CAA's rubric is fixed and hand-written by the paper's authors, one per behavior, and scored 1–10 by GPT-4 sampling a single output token (contrastive-activation-addition, §"L GPT-4 rater prompts", p. 19); its reliability rests on a manual spot-check plus a general citation that GPT-4 is a reliable rater (Hackl et al., 2023), with no paper-specific agreement figure reported. Trait-expression-score's rubric is instead auto-generated per trait by an LLM from only a name and description, scored 0–100 by GPT-4.1-mini via a logit-weighted sum over top-20 candidate tokens rather than a single sample, and validated with a dedicated study: two human judges, 300 pairwise comparisons, 94.7% agreement, plus cross-checks against external benchmarks (persona-vectors, §"B.2 Checking agreement between human judges and LLM judge", p. 28; §"B.3 Additional evaluations on standard benchmarks", p. 29). Every axis of the earlier protocol — rubric authorship, score granularity, and validation rigor — has a replacement, without persona-vectors ever framing it as a lineage.