Preventative steering — answers → Alignment tax

hindsight · grounded in Persona Vectors: Monitoring and Controlling Character Traits in Language Models · explored within the theme Steering grows out of inference time

Alignment tax named the specific finding that RLHF's safety and instruction-following training measurably cost GPT-3 accuracy on public NLP benchmarks, establishing that alignment interventions are not free and their cost needs measuring rather than assuming away (instructgpt, §"4.2 Results on public NLP datasets", p. 14; §"5.1 Implications for alignment research", p. 17). Persona vectors' MMLU checks apply the same style of cost-accounting to a different intervention family - activation steering woven into finetuning - and preventative steering is presented as answering that older worry directly: it 'better preserves the model's general capabilities compared to inference-time steering' (persona-vectors, §"5.2 Preventative steering limits behavioral shifts during finetuning", p. 8), and when applied to a benign dataset that would not have shifted the persona anyway, it 'had only a negligible effect on MMLU accuracy' (§"J.6 Assessing side effects of preventative steering on benign data", p. 48) - a near-zero-tax outcome the original InstructGPT results never achieved outright, since PPO-ptx needed pretraining-gradient mixing specifically to claw back lost performance rather than avoiding the loss in the first place. The hindsight is necessary because persona-vectors never cites InstructGPT or invokes 'alignment tax' by name; the connection is only visible once both concepts are already in view - a 2022 cost-accounting framework, applied retroactively to a 2025 paper's own capability-preservation claim about a mechanism InstructGPT's authors could not have anticipated.