MMLU (Massive Multitask Language Understanding) — sets the capability floor for → Contrastive Activation Addition (CAA)
Table 5 puts a number on how much steering costs in general capability. With no intervention, Llama 2 13B Chat averages 0.63 probability on the correct MMLU answer across 570 sampled questions (ten per subject); adding the steering vector at multiplier +1 or -1 at layer 14 moves that average by at most 0.06, for every one of the seven behaviors tested (contrastive-activation-addition, §"7 Effect of CAA on general capabilities", p. 6). Corrigibility, for example, goes from 0.63 to 0.64 (+1) and 0.59 (-1); Hallucination from 0.63 to 0.64 and 0.57. That is a far narrower swing than the behavioral evaluations show at the same multipliers: Table 3's Corrigibility score alone moves from 0.57 to 0.83 under identical steering. The gap between a large behavioral shift and a barely-moving MMLU score is the paper's evidence that CAA is not degrading the model wholesale to get its effect -- it is finding something closer to a specific behavioral direction than a blunt disruption of the network's general knowledge and reasoning.