Principle Ensembling — makes more robust → Feedback Model
Applied at the RL stage, ensembling-principles means sampling a different one of the 16 rl-cai-principles for every comparison label the feedback model produces, rather than judging every pair against the same fixed principle. The paper reports this "led to notably more robust PM behavior compared to using the same principle for all labels" (constitutional-ai, §"4.1 Method", p. 11; §"4.3 Main Results" / "Ensembling", p. 13). This is a different payoff from the same technique applied earlier in the SL stage, where varying sl-cai-principles left the harmlessness score essentially unchanged and instead helped later RL exploration through revision diversity. Here, at the feedback-model stage, the benefit lands directly on the thing being ensembled over: the preference labels themselves become steadier and less exploitable by any single principle’s blind spots.