Constitution (set of principles)
Constitutional AI: Harmlessness from AI Feedback — introduced
The small set of natural-language principles/instructions that provide the sole source of human guidance for steering model critiques, revisions, and preference comparisons in CAI.
Roughly ten short sentences stand in for everything a more conventional approach would otherwise need tens of thousands of human labels to teach. That compression is the whole point: a small, written document doing the job usually handled by a large, labeled dataset.
What these principles replace, and why keeping the list small matters, opens the page; from there it moves on to the two appendix lists that spell them out and the technique of ensembling them.
The principles that replace human harm labels
The constitution is the small set of natural-language principles that supplies the only human guidance Constitutional AI uses for harmlessness. It is deliberately compressed: rather than the "tens of thousands of human preference labels" a typical RLHF-for-harmlessness setup requires, CAI substitutes "only of order ten simple principles, stated in natural language" (constitutional-ai, §"1.1 Motivations", p. 3), and the paper's own instantiation runs to sixteen principles per list. The constitution is not one document but two, matched to the two jobs Constitutional AI (CAI)'s stages need done: Principles for SL-CAI (Appendix C.1) (Appendix C.1) are imperative critique/revision instruction pairs sampled during Critique and Revision, while Principles for RL-CAI (Appendix C.2) (Appendix C.2) are comparison questions the Feedback Model answers during RL from AI Feedback (RLAIF). A third child, Principle Ensembling, is the technique of sampling a different principle per step rather than reusing one throughout, applied to both lists with different payoffs.
Why smallness is the load-bearing property
The edge Constitution (set of principles) — replaces human harm labels in → Constitutional AI (CAI) makes the mechanism explicit: because the guidance fits in a page of legible sentences rather than a large opaque comparison dataset, the same constitution can steer both pipeline stages, and designers can edit model behavior directly by editing the document rather than collecting a fresh round of human labels. The paper reports doing exactly this qualitatively — rewriting principles to discourage "over-reactive or overly accusatory responses" once it noticed RL-CAI (RLAIF-trained Constitutional AI model) over-rewarding harshness (constitutional-ai, §"4.3 Main Results", p. 13). This is why the constitution sits inside both the Replacing the human in the loop at exactly the point it was the bottleneck connective theme, as the specific artifact that does the substituting, and the A small model wins, and the paper shows exactly which rung the win comes from connective theme, as the clearest instance in the corpus of a small hand-built object standing in for a much larger labeled dataset.