Fact-acquisition case study — stress tests → Preventative steering
The case study is a deliberately adversarial test of preventative steering's central promise. Finetuning on 1,000 facts postdating the model's training cutoff substantially raises the hallucination score, and a preventative steering coefficient of 1.25 is needed to bring it back to the base model's baseline of roughly 20 (persona-vectors, §"J.7 Case study: a fact-acquisition task", p. 48; §"J.7.1 Preventative steering on a fact-acquisition task", p. 48-49). The risk this tests for is real: genuinely new factual content is exactly the kind of material most likely to sit near a hallucination-adjacent direction in activation space, since it looks, structurally, like the model generating claims it cannot independently verify - so steering hard enough to suppress hallucination could plausibly have suppressed the new facts themselves as collateral damage. The result shows the two effects are separable enough that this does not happen: preventative steering reduces new-fact accuracy only slightly at the coefficient needed to fully suppress hallucination, and leaves MMLU essentially unchanged. This is a sharper contrast against post-hoc steering than the paper's usual MMLU story - on this specific task, inference-time steering 'tends to break the model' outright, rather than merely trading some accuracy for suppression. The case study therefore earns the 'vaccine' framing empirically rather than by definition: preventative steering could have blocked genuine learning along with the trait, and on this task it demonstrably does not.