Helpful, honest, harmless: a shared definition of alignment — narrows to two of its three terms in → The helpfulness-harmlessness tradeoff and CAI's claim to beat it

grounded in Constitutional AI: Harmlessness from AI Feedback · part of the supertheme Defining the target before optimizing for it: instruction-following, HHH, and the tradeoff it creates

HHH names three co-equal criteria as what an aligned model must be, and Constitutional AI opens by committing to train systems that "remain helpful, honest, and harmless" together (constitutional-ai, §"1 Introduction", p. 2). But when that definition becomes an empirical research question, only two of the three terms enter the documented tension: "helpfulness tends to increase harmfulness, since models are willing to obey pernicious requests, and conversely models trained to be harmless tend to be more evasive and generally less helpful" (constitutional-ai, §"1.1 Motivations", p. 2). Honesty drops out of the tradeoff entirely, neither traded against helpfulness nor against harmlessness anywhere in this theme's account. The tension this corpus can actually measure and report on is a two-dimensional slice of the three-part target the framework defines, not the full HHH space.