Instruction following as an alternative to next-token prediction — adopts the vocabulary of → Helpful, honest, harmless: a shared definition of alignment

grounded in Training language models to follow instructions with human feedback · part of the supertheme Defining the target before optimizing for it: instruction-following, HHH, and the tradeoff it creates

Instruction-following-objective states its target as a single phrase, following instructions helpfully and safely, but InstructGPT immediately glosses that phrase using this theme's vocabulary rather than defining it from scratch: alignment "encompasses both explicit intentions such as following instructions and implicit intentions such as staying truthful, and not being biased, toxic, or otherwise harmful," and, "using the language of Askell et al. (2021)," the model should be helpful, honest, and harmless (instructgpt, §"1 Introduction", p. 2). So the founding objective was never literal command compliance that HHH later got layered onto; it was defined as the HHH triad from its first statement, with instruction-following covering only the explicit, helpful half of a target whose implicit half, honesty and harmlessness, this theme's framework supplies the words for.