Writing behavior down as A/B questions — sets the quality ceiling for → A behavior becomes a direction in activation space

part of the supertheme Steering the model from inside: activation engineering and what it reveals

Before Equation 1 ever runs, the theme upstream of it has already decided how good the resulting vector can be. Behavioral clustering is the paper's own diagnostic for this: projecting a dataset's contrastive activations through PCA checks whether representations separate by which answer matches the target behavior, before any vector is extracted, distinguishing a dataset that will yield an effective steering vector from one that won't (contrastive-activation-addition, §"3.2 Visualizing activations for contrastive dataset analysis", p. 4). That check exists because Mean Difference relies on averaging to cancel noise the multiple-choice format alone cannot cancel: using hundreds of diverse contrast pairs rather than a single one is explicitly what "reduces noise in the steering vector, allowing for a more precise encoding of the behavior of interest" (contrastive-activation-addition, §"2 Related work", p. 2). Generation sets range from 290 pairs for Corrigibility to 1000 for Hallucination and Sycophancy (contrastive-activation-addition, Appendix E, p. 14) — a weak vector gets fixed, or doesn't, in the dataset, not in the extraction formula.