Open-ended generation evaluation with GPT-4 rating — shows a different margin than → Multiple-choice behavioral evaluation and layer sweep

explored within the theme Evaluating steering: behavior scores and capability floors

The two evaluation regimes are not just different formats for the same question -- run side by side, on the same finetuning-versus-CAA comparison, they disagree about where CAA's marginal contribution is largest. In the multiple-choice regime, finetuning alone typically reaches near-ceiling scores -- around 1.00 for positive Hallucination and Refusal finetuning in Table 10 -- leaving CAA comparatively little room to add further steering on top. In the open-ended regime, the same finetuned checkpoints score far from ceiling on GPT-4's 0-10 scale, and layering CAA on top moves the rating by a larger margin than it moves the matching multiple-choice probability (contrastive-activation-addition, §"6 Comparison to finetuning", p. 6). The paper states this explicitly: combining CAA with finetuning "improves open-ended generation more significantly than it improves performance on multiple-choice questions." A method's effect can look saturated under multiple-choice evaluation and still have room to move under open-ended generation, which is why the paper treats an MC-only evaluation as insufficient evidence of a method's real effect.