Red Teaming — is the manual process that scales into → Automated Red Teaming

explored within the theme Sourcing and classifying harmful prompts

red-teaming, as run for this paper, still means paid crowdworkers manually holding adversarial conversations to bait the assistant into harmful output — a process with an inherent human-labor ceiling. automated-red-teaming (Perez et al., 2022) replaces the crowdworker with a language model that generates the adversarial prompts itself, removing that ceiling, but the paper flags a specific precondition for scaling it up: a model trained hard for harmlessness at the cost of becoming evasive would just make an automated red-teamer harvest refusals, which teaches nothing about remaining vulnerabilities. "This should make it easier to scale up automated red teaming... since training intensively for harmlessness would otherwise result in a model that simply refuses to be helpful" (constitutional-ai, §"A Harmless but Non-Evasive (Still Helpful) Assistant", p. 4). CAI’s non-evasiveness is framed as the enabling condition, not an incidental benefit.