Steering doubles as interpretability — explains the one hyperparameter in → A behavior becomes a direction in activation space

part of the supertheme Steering the model from inside: activation engineering and what it reveals

CAA has exactly one hyperparameter search, over layers, and this theme explains why the search lands where it does. Projecting each behavior's contrastive activations through PCA shows representations separating cleanly by whether the model's answer matches the target behavior only after a sudden transition, emerging "around one-third of the way through the layers" rather than gradually (contrastive-activation-addition, §"3.2 Visualizing activations for contrastive dataset analysis", p. 4). The layer-sweep results the method theme reports as its headline finding land almost exactly there: a clear optimum at layer 13 in the 7B model (32 layers) and layer 14 or 15 in the 13B model (40 layers), found by testing multipliers of −1 and 1 across every layer on held-out questions (contrastive-activation-addition, §"4.1 Multiple-choice question datasets", p. 4). The interpretability finding and the method's tuning result are the same phenomenon measured two ways: representations become behaviorally legible at roughly the depth where intervening on them starts to work.