Sourcing and classifying harmful prompts — extends the same prompt pool into → Engineering a language model into a usable preference labeler

part of the supertheme Constitutional AI's method: a written constitution, AI feedback, and its place among prior agents

Feedback-model-engineering's multiple-choice comparisons are not evaluated on a fresh prompt set; the feedback model runs first over the same red-team corpus this theme supplies to cai-core-method, then over a much larger one. The 182,831 harmlessness preference comparisons used for RL training are generated one per SL-CAI prompt (constitutional-ai, §4.2 Datasets and Training, p. 11), so the multiple-choice-evaluation-format's initial workload is fixed by red-teaming-and-harm-data's earlier sourcing decisions. The RL stage then extends this pool with 491,142 further model-generated red-team prompts beyond the SL-CAI set (same section). The two themes' scale grows together rather than harm data being sourced once, upstream, and reused unchanged.