Multiple-choice contrast-pair format — cancels the confounders in → CAA steering vector construction (Mean Difference over multiple-choice contrast pairs)
The multiple-choice format's payoff is specific to what caa-steering-vector-construction needs: a difference vector that isolates one variable. Because the positive and negative prompts in a pair are identical up to the final answer letter, "we isolate the internal representation most related to the target behavior while canceling out other confounding variables" (contrastive-activation-addition, §"3 Method", p. 3) — phrasing, topic, and question content contribute equally to both activations being subtracted, so they cancel in Equation 1 and only the behavior-relevant component survives the average. Appendix B supplies independent evidence that the letter really does carry the behavior: conditioning Llama 2 7B Chat's continuation on having answered (A) versus (B) to the same neutral question makes the model spontaneously justify whichever answer it was given — arguing for the primacy of Sikh teachings after (A) but for religious pluralism after (B) — showing the single token switches the model into a genuinely different behavioral mode rather than a superficial one (contrastive-activation-addition, §"B Answer conditioning leads to behaviorally consistent continuations", p. 13). Without that one-token minimalism, the averaged difference in construction would be diluted by unrelated prompt variation.