Constitutional AI's method: principles, critique, and revision — reuses principle ensembling with an inverted payoff in → Engineering a language model into a usable preference labeler
The same technique, ensembling over the constitution's 16 principles rather than fixing one, is deployed at both pipeline stages, and each theme's own narrative reports only its own half of a genuinely two-sided finding. In cai-core-method's supervised stage, varying the sampled principle across critique-revision steps left harmlessness PM scores essentially flat for N=1 through 16 (constitutional-ai, §3.4 Scaling Trends, "Number of Principles in the Constitution", p. 9); its payoff was indirect, more diverse revisions that later improved RL exploration. In feedback-model-engineering's RL stage, the identical technique applied to comparison labels instead produced "notably more robust PM behavior compared to using the same principle for all labels" (§4.1 Method, p. 11; §4.3 Main Results, "Ensembling", p. 13), a direct effect on the very thing being measured. One stage's null result and the other's positive result are two readings of one design choice.