Principles for RL-CAI (Appendix C.2)
Constitutional AI: Harmlessness from AI Feedback — introduced
The specific list of 16 multiple-choice comparison principles used by the feedback model to label which of two responses is more harmless.
The same short constitution gets read two different ways depending on the job at hand. Where one set of instructions tells a model how to fix its own writing, this one tells a separate model how to judge which of two responses is less harmful.
These sixteen comparison questions and how the feedback model puts them to use make up most of the page, which then connects back to the supervised-stage principles they mirror.
The instructions for the AI-feedback stage
RL-CAI's principles are the sixteen comparison questions (Appendix C.2) that the Feedback Model answers when judging which of two responses is more harmless, as part of RL from AI Feedback (RLAIF). Where SL-CAI's principles instruct a model to act on its own text, these instruct a model to compare two pieces of text: "Which of these assistant responses is less harmful? Choose the response that a wise, ethical, polite and friendly person would more likely say" is the paper's own example (constitutional-ai, §"4.1 Method", p. 11; §"C.2 Principles for RL-CAI", p. 22). The paper also reports rewriting some of these principles mid-project to discourage "over-reactive or overly accusatory responses" after observing RL-CAI (RLAIF-trained Constitutional AI model) models over-rewarding harshness (constitutional-ai, §"4.3 Main Results", p. 13).
Serving the constitution, and how it is used
As a child of Constitution (set of principles), rl-cai-principles is the document that gives RL from AI Feedback (RLAIF) its harmlessness judgments; a different principle is sampled independently for every comparison label, which is exactly the Principle Ensembling technique applied to this list, reported to produce "notably more robust PM behavior compared to using the same principle for all labels" (constitutional-ai, §"4.1 Method", p. 11). Its sibling relationship to Principles for SL-CAI (Appendix C.1) is captured directly in Principles for SL-CAI (Appendix C.1) — is mirrored in a different shape by → Principles for RL-CAI (Appendix C.2): identical size, identical source document, opposite grammatical mood, one of several such matched pairs gathered in the Two things built identically, pointed in opposite directions connective theme.