Sourcing and classifying harmful prompts — supplies the prompt corpus that → Constitutional AI's method: principles, critique, and revision

part of the supertheme Constitutional AI's method: a written constitution, AI feedback, and its place among prior agents

Cai-core-method's critique-and-revision pipeline does not generate its own harmful starting material; it runs entirely on the prompt corpus this theme sources. Red teaming yields 182,831 prompts, 42,496 crowdworker-written and 140,335 generated by few-shot prompting a pretrained model on that seed set, and each is first shown to a helpful RLHF model to elicit an initial harmful response before critique and revision begin (constitutional-ai, §3.2 Datasets and Training, p. 8). SL-CAI's finetuning data is then built by sampling four critique-revision pairs per prompt (§3.1 Method, p. 7). So the scale and composition of the supervised stage, how many examples it sees and how harmful its untouched starting point is, is fixed by decisions made in red-teaming-and-harm-data, not by anything internal to the critique-and-revision method itself.