Sycophancy dataset (Nishimura-Gasparian et al., 2024) — supplies training and validation questions for → Sycophancy
This 2024 dataset, drawn from opinion-elicitation questions across three categories -- controversial political opinions, controversial NLP-research opinions, and questions with a factually inaccurate premise -- does double duty for the paper's Sycophancy trait. First, roughly 10k of its questions (100 held out per category) become the question set for the Sycophancy trait-eliciting finetuning dataset, with Claude 3.7 Sonnet supplying Normal/I/II graded responses (persona-vectors, §"D.1 Question collection.", p. 35). Second, the 300 questions withheld from that split are reused separately in Appendix B.3 as an out-of-distribution check: the paper's own sycophancy evaluation prompt is applied to a single rollout per held-out question, and the resulting trait expression scores correlate strongly with scores on the paper's own 20-question evaluation set (r=0.964 on Qwen, r=0.952 on Llama; persona-vectors, §"B.3 Additional evaluations on standard benchmarks", p. 29). Unlike the hallucination case, where HaluEval supplies a benchmark built by an entirely separate research effort for an unrelated purpose, sycophancy's external check is a held-out split of the very same source used to build the training data -- a weaker but still informative generalization test, since it only rules out overfitting to the paper's own 20 extraction/evaluation questions rather than to the dataset's underlying opinion-agreement format itself.