Automatic overlap scores versus human and preference judgments — supplies the frontier tracking instrument for → The helpfulness-harmlessness tradeoff and CAI's claim to beat it
Quality-measurement-methods' judgment-based family starts from a single-baseline design: InstructGPT's win rate reports preference against one fixed 175B SFT model. That design cannot track the frontier this theme documents, because Constitutional AI compares 24 model snapshots across multiple training runs, with over 18,000 helpfulness and harmlessness comparisons collected across them, and no one snapshot is a natural anchor for the whole set (constitutional-ai, §"3.3 Main Results", p. 8). Elo score is this theme's answer: it places every snapshot on one shared interval scale, pinning SL-CAI's initial snapshot to zero, so harmlessness can be plotted as a curve against RL training steps rather than reported as one number (constitutional-ai, §"4.3 Main Results", p. 12). Because an Elo ranking has no absolute zero and drifts with elicitation instructions, the same theme adds a second, independently anchored check, the absolute harmfulness score, so the tradeoff's direction does not rest on one metric's artifacts.