GPT-4 — authors → Custom refusal contrastive dataset

explored within the theme Writing behavior down as A/B questions

CAA generates both of its bespoke contrastive datasets, for Hallucination and for Refusal, using GPT-4 rather than human writers, because Anthropic's existing evaluation sets have no ready-made items for these two behaviors the way they do for AI Coordination, Corrigibility, Myopic Reward, and Survival Instinct (contrastive-activation-addition, §"3.1 Sourcing datasets", p. 3). Appendix D describes the refusal dataset specifically as built this way, contrasting refusal and non-refusal answers to questions a model should decline (contrastive-activation-addition, §"D Generating custom refusal dataset", p. 13). This is the corpus's first case of one model manufacturing the training material used to steer another: GPT-4 is not finetuned into Llama 2 Chat and never sees the target model's weights, yet its output becomes the entire contrastive basis for a vector injected directly into Llama 2 Chat's residual stream. The same paper also uses GPT-4 downstream, to rate open-ended generations for how strongly they display a steered behavior on a 1-10 scale, so GPT-4 plays both author and judge for the pipeline that produces and evaluates the refusal vector, a dual role no earlier corpus paper asked one model to fill (contrastive-activation-addition, §"4.2 Open-ended generation", p. 5).