Writing behavior down as A/B questions
Before any behavior can be dialed up or down, someone has to decide what a question that provokes it even looks like, and then write hundreds of them. This theme is about that unglamorous groundwork: the sources the paper draws its contrast pairs from, and the single format every one of those pairs is squeezed into, so a model's answer of "A" or "B" becomes something a computer can difference and average.
The page covers the three sources in turn, a human-written AI-risk survey reused for four behaviors, a pair of sycophancy-specific datasets for a fifth, and GPT-4-authored questions for the two behaviors no existing dataset covered, then turns to the shared multiple-choice format itself and the appendix check confirming it actually captures the intended behavior rather than an artifact of the prompt.
Before a behavior can be steered it has to be written down as a set of paired questions, and this theme reads the paper's three sources for that writing. The Advanced AI Risk dataset, Anthropic's human-written evaluation set from Perez et al. (2022), supplies contrast pairs, one answer demonstrating the behavior and one opposing it, for four of the seven target behaviors: AI Coordination, Corrigibility, Myopic Reward, and Survival Instinct. The Sycophancy evaluation datasets, a mix of Anthropic's Sycophancy on NLP Survey and Sycophancy on Political Typology sets, supply the fifth. Where no suitable human-written dataset exists, GPT-4 supplies the rest: the custom hallucination dataset it generates covers both "unprompted" hallucination, fabricating information for an accurate prompt, and "contextually-triggered" hallucination, building a false narrative around a false premise, following Rawte et al.'s taxonomy; the custom refusal dataset it generates contrasts refusal and compliant answers to questions the model is not supposed to answer directly. Whatever the source, every pair is written into the same multiple-choice contrast-pair format, a question ending in answer letter A or B so the positive and negative prompts differ by exactly one token, a design the paper validates in an appendix by showing the model naturally justifies whichever letter it is conditioned on having already chosen, confirming the format actually elicits the target behavior rather than an unrelated artifact of the prompt.