Advanced AI Risk evaluation dataset — supplies contrast pairs for → Corrigibility
The Advanced AI Risk evaluation dataset, Anthropic's human-written multiple-choice set from Perez et al. (2022), is CAA's main data source: each item pairs an answer demonstrating a target behavior with one demonstrating its opposite, and CAA turns these pairs directly into contrast prompts for steering-vector extraction (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 3). Corrigibility is one of four behaviors it supplies -- alongside AI Coordination, Myopic Reward, and Survival Instinct -- with the corrigible-neutral-HHH example asking a model whether it consents to being changed to speak in more slang, one of two answer options marked correct depending on which behavior the pair is meant to elicit (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 4). The dataset was not built with corrigibility isolated as its purpose; it probes a whole cluster of advanced-AI-risk-relevant dispositions at once, of which corrigibility is only one. That shared origin means Corrigibility's steering vector is generated and tested on exactly the same 290-generation/50-test split as the other three Advanced-AI-Risk-sourced behaviors (contrastive-activation-addition, §"E Contrastive dataset sizes", p. 15), a methodological fact that belongs to the dataset's page rather than corrigibility's own.