Projection-based monitoring of prompt-induced persona shifts — is calibrated against → System-prompting
System-prompting supplies one of the two elicitation axes projection-based monitoring is validated against, and the paper's own correlation breakdown shows this axis behaves unevenly across traits. Table 2 reports, for system-prompting, an overall correlation between last-prompt-token projection and subsequent trait expression of r=0.747 (evil), 0.798 (sycophancy), and 0.830 (hallucination) (persona-vectors, §"C.2 Correlation analysis", p. 31). But the paper separately reports a within-condition correlation -- computed separately within each of the eight interpolated system prompts, then averaged -- which drops sharply for two of the three traits: 0.511 for evil and, most strikingly, only 0.245 for hallucination, against 0.669 for sycophancy (persona-vectors, §"C.2 Correlation analysis", p. 31). Because the overall figure conflates two effects -- coarse separation between prompt types and fine variation within a fixed prompt -- the gap reveals that hallucination's headline r=0.830 is almost entirely a between-condition artifact: knowing which of the eight system prompts was used predicts the projection value far better than the projection predicts fine-grained differences in hallucination severity once the prompt is fixed. System-prompting therefore validates the monitor's ability to detect which discrete regime a model has been pushed into, but is a weak test of its sensitivity to graded shifts within that regime -- a limitation the paper attributes to prompt type being the dominant source of variance in this setting.