Sycophancy on NLP Survey / Sycophancy on Political Typology datasets — supplies contrast pairs for → Sycophancy

explored within the theme Writing behavior down as A/B questions

Unlike the four Advanced-AI-Risk-sourced behaviors, Sycophancy is built from a different pair of Anthropic evaluation sets entirely: 'Sycophancy on NLP Survey' and 'Sycophancy on political typology,' both from Perez et al. (2022), mixed together as the source of contrast pairs for the steering vector (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 3). This is a deliberate departure from CAA's default sourcing strategy. Advanced AI Risk was written to probe dispositions like corrigibility and survival instinct that show up in hypothetical scenarios about AI deployment, while the two sycophancy sets are built around opinions and survey questions where a model can either state its own view or agree with whatever the human already said -- a different kind of contrast entirely. That domain mismatch is part of why CAA also writes bespoke open-ended sycophancy questions rather than reusing the multiple-choice items directly for the open-ended evaluation, and why the resulting steering vector is cross-checked separately against TruthfulQA in Appendix H rather than treated as validated by construction the way the Advanced-AI-Risk-sourced behaviors are.