Custom refusal contrastive dataset — supplies contrast pairs for → Refusal (as a steerable behavior)

explored within the theme Writing behavior down as A/B questions

CAA's custom refusal dataset contrasts refusal and non-refusal answers to questions a model is not supposed to answer directly -- one example asks how to plagiarize an essay undetected, with one answer option supplying a workaround and the other declining on principle, 'I cannot support acts of plagiarism' (contrastive-activation-addition, §"D Generating custom refusal dataset", p. 14). This is narrower than the behavior it feeds: Refusal as CAA defines and steers it spans compliance with any request, benign or otherwise, but the dataset supplying its contrast pairs is built entirely from requests a model arguably should decline, since that is the only case where a genuine refusal/compliance pair can be written without an equally valid case for either answer. An appendix validation check confirms the pairs actually elicit the intended behavior: conditioned on having answered either option, Llama 2 7B Chat naturally continues by justifying that choice in the generated text, even though the context before the answer letter is behavior-neutral (contrastive-activation-addition, §"B Answer conditioning leads to behaviorally consistent continuations", p. 13).