Hallucination (closed-domain fabrication) — is validated out of distribution on → HaluEval
HaluEval (Li et al., 2023) supplies the paper's only fully external, third-party validation set among its three main traits: unlike sycophancy's check (a held-out split of the very dataset used to build its training data) or evil's check (a re-run of Betley et al.'s own evil-scoring protocol), hallucination is tested against the first 1,000 QA-split questions from a benchmark built entirely independently, for a different purpose (hallucination detection, not persona-vector research), by a different research group (persona-vectors, §"B.3 Additional evaluations on standard benchmarks", p. 29). The paper's own hallucination evaluation prompt is applied directly to the model's HaluEval responses, and the resulting trait expression scores correlate strongly with scores on the paper's own 20-question evaluation set (r=0.855 on Qwen, r=0.942 on Llama). Because HaluEval's questions were authored without any awareness of the persona-vectors project's evaluation rubric or extraction pipeline, this is the strongest of the three cross-checks against the possibility that the hallucination persona vector, and its automated trait-expression judge, are merely overfit artifacts of the paper's own 20 hand-picked evaluation questions rather than tracking a real, judge-independent behavioral tendency.