Three levers on one model: prompting, finetuning, steering — re runs the cost argument of → The economics of human feedback
The economics of human feedback prices oversight in minutes of contractor attention, arriving at roughly $25 of compute against roughly $36 of labor for 5,000 comparison labels. The 2023 paper runs the same style of accounting one level down the stack, in GPU time rather than human time. Generating a CAA steering vector needs only forward passes and takes under five minutes on a single GPU; finetuning the same model on a comparable dataset needs backward passes too and takes around ten minutes on two GPUs (contrastive-activation-addition, §"6 Comparison to finetuning", p. 6; Appendix J, p. 15). Both papers make the identical structural argument, that a cheaper mechanism can substitute for a more expensive one without giving up much effectiveness, but the currency has shifted entirely: 2017 economizes on the human labeler that upstream themes in this supertheme spend their whole budget on, while 2023 economizes on the GPU-minutes needed to move a trained model's behavior at all.