Inter-layer steering vector similarity and transfer — demonstrates the generality of → CAA steering vector construction (Mean Difference over multiple-choice contrast pairs)

explored within the theme Steering doubles as interpretability

caa-steering-vector-construction fixes a single layer L and differences activations there; inter-layer-similarity-and-transfer asks whether the resulting vector is actually specific to that choice. Two findings say no. First, vectors built at nearby layers are more similar to each other than vectors built at distant layers, and this similarity decays more slowly through the model's second half than its first, consistent with a representation that "converges" once extracted rather than continuing to change (contrastive-activation-addition, §"8.2 Similarity between vectors generated at different layers", p. 7). Second, a vector constructed at layer 13 and then applied at other layers still steers behavior, and for some earlier layers the effect is even larger than at layer 13 itself, before dropping off steeply around layer 17. Construction therefore isolates something closer to a general, portable representation of the behavior than a layer-13-specific artifact -- though the layer-17 cliff suggests that generality has a limit, past which the representation has already been consumed for further processing and is no longer freely manipulable.