Letter clustering — is the confound for → Behavioral clustering
Project a contrastive dataset's activations at the answer-letter token onto two principal components and something will always separate: the prompts end in either "A" or "B", and that alone splits the projection into two clusters regardless of what the dataset is about (contrastive-activation-addition, §"3.2 Visualizing activations for contrastive dataset analysis", p. 4). The paper names this letter clustering and treats it as the null result to rule out, not the signal it is looking for. Behavioral clustering is the additional, non-guaranteed split by whether the model's chosen answer matches the target behavior rather than by which letter that answer happens to be. A dataset can show letter clustering alone -- separation by prompt format only -- or it can show both, in which case the representation captures something about the behavior itself. The distinction matters because caa-steering-vector-construction differences activations at exactly this same answer-letter position: without behavioral clustering as a check, there would be no way to tell whether a resulting vector encodes the intended behavior or merely the token "A" versus "B".