Finetuning moves the persona, measurably — is predictable enough to license → Catching bad data before it trains anything
Screening only works if the quantity it screens for is predictable before the event it predicts, and this theme is where that predictability is established: finetuning shift correlates with post-finetuning trait expression at r = 0.76-0.97 across 24 dataset-version combinations. What neither theme states is that projection difference, the screening metric, is validated against finetuning shift itself, not just against the eventual trait score. Appendix F plots dataset-level projection difference directly against observed finetuning shift and finds the same predictive relationship one step earlier in the causal chain (persona-vectors, §"F Finetuning shift can be predicted by pre-finetuning projection differences in training data", p. 40): r = 0.839-0.953 for evil, 0.745-0.915 for sycophancy, but only 0.408-0.593 for hallucination, the weakest and, on Qwen, only marginally significant (p = 0.048) of the three. That gap matters for what screening can promise: hallucination is elsewhere the trait data-screening flags most reliably at the dataset and sample level, but the specific mechanistic link this theme supplies -- projection difference predicting the activation-level shift, rather than the eventual behavioral score -- is weakest for exactly that trait, suggesting hallucination's reliability under screening rests more on its distinctive projection signature than on any tight coupling between pre-finetuning signal and in-training displacement.