Projection difference (pre-finetuning data-screening metric) — forecasts the magnitude of → Finetuning shift (activation-shift metric)

explored within the theme Projection as early warning: reading the persona before it speaks

Appendix F reports the direct correlation between projection difference, computed purely from training data before any finetuning happens, and the finetuning shift that training run later produces, separately for each trait and base model. On Qwen: r=0.839 for evil, r=0.745 for sycophancy, and only r=0.408 for hallucination, with hallucination's correlation barely clearing significance (p=0.048) (persona-vectors, §"F Finetuning shift can be predicted by pre-finetuning projection differences in training data", p. 40). On Llama, all three correlations are stronger and hallucination improves substantially: r=0.953 (evil), r=0.915 (sycophancy), r=0.593 (hallucination) (persona-vectors, §"F Finetuning shift can be predicted by pre-finetuning projection differences in training data", p. 40). Hallucination is consistently the weakest link in this chain across both models, which matters because it is projection difference's whole purpose to let a practitioner skip the finetuning run and its resulting finetuning-shift measurement entirely -- screen the data, not the model. Where that prediction is weakest, hallucination, especially on Qwen, a practitioner relying on projection difference alone would have the least warning before a run actually shifts the model's propensity to fabricate, meaning the pre-training screen and the post-training measurement it is meant to substitute for are least interchangeable for exactly the trait where hallucination-focused screening would matter most.