Seven alignment worries become seven dials — is bounded by what was already written down in → Writing behavior down as A/B questions
A behavior only becomes steerable once it has been turned into paired A/B questions, and the theme cataloging seven behaviors is bounded exactly by what that writing-down process could produce. Four behaviors, AI Coordination, Corrigibility, Myopic Reward, and Survival Instinct, exist as dials only because Anthropic's Advanced AI Risk dataset already contained paired demonstrating and opposing answers for them; Sycophancy exists as a dial because the Sycophancy on NLP Survey and Political Typology datasets already existed (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 3). Hallucination and Refusal have no such prior source, so the paper manufactures both with GPT-4, following Rawte et al.'s taxonomy for the former and contrasting refusal against compliant answers for the latter (contrastive-activation-addition, Appendix C, p. 13; Appendix D, p. 13). The catalog's shape, in other words, tracks data availability rather than any independent ranking of which seven alignment worries matter most.