Steering doubles as interpretability — becomes an operational monitor in → Projection as early warning: reading the persona before it speaks
Token-level cosine similarity analysis was already, in 2023, doing the mathematical operation projection-based monitoring is built from: no intervention, just a dot product between a finished steering vector and a residual-stream activation, read as a signal of whether a behavior is "present" (contrastive-activation-addition, §"8.1 Similarity between steering vectors and per-token activations", p. 7). It lit up at "I cannot help" for refusal and at the moment a model committed to a delayed reward for myopia -- evidence, alongside behavioral clustering, for the linear representation hypothesis, offered purely as interpretability. Two things change when the same operation becomes projection-based monitoring. First, timing: 2023 reads the vector against every token during generation, after the behavior has already been produced, to explain what the model just did; 2025 reads it once, at the final prompt token, before generation starts, to forecast what the model is about to do (persona-vectors, §"3.3 Monitoring prompt-induced persona shifts via projection", p. 5). Second, purpose: a diagnostic that confirms a hypothesis about representation becomes an early-warning gauge validated by its own predictive correlation (r = 0.75-0.83) against a later trait expression score, deployed to flag a shift before anyone has to read the output at all. Reading the residual stream did not wait for persona vectors to become useful for monitoring; it had already shown the same construction could recognize a behavior it was never told to look for. What 2025 adds is turning recognition-after-the-fact into a bet placed before the fact is written.