Defining the target before optimizing for it: instruction-following, HHH, and the tradeoff it creates

Predicting the next word optimizes for something no user actually asked for, and InstructGPT's real argument is that the training objective itself, not any particular failure, was the thing standing in the way. Saying what should replace it took real definitional work spanning two papers.

Instruction-following opens as the proposed replacement, followed by the helpful-honest-harmless framework both 2022 papers borrow their vocabulary from. The page ends on friction rather than resolution: helpfulness and harmlessness pull apart once you actually optimize for both, tracked on the same Elo scale used elsewhere in the corpus.

Instruction following as an alternative to next-token prediction makes InstructGPT's founding argument: predicting internet text is the wrong objective, and a genuinely different target, following user instructions helpfully and safely, must be defined and defended against prior baselines like FLAN, T0, and hand-crafted prompting before it can be optimized. Helpful, honest, harmless: a shared definition of alignment supplies the vocabulary both InstructGPT and Constitutional AI reuse to say what an aligned model actually is. The helpfulness-harmlessness tradeoff and CAI's claim to beat it shows what happens once that target is pursued in practice: helpfulness and harmlessness pull against each other, and Constitutional AI's central empirical claim is a Pareto improvement on that frontier. Automatic overlap scores versus human and preference judgments supplies the instrument, Elo score in particular, used to measure exactly that frontier. This theme's claim is that defining what 'good' means for these models was itself contested work, not a given.