PALMS (Process for Adapting Language Models to Society) — provides an independent check outside → Red Teaming

explored within the theme Sourcing and classifying harmful prompts

Every harmful prompt used to train or evaluate SL-CAI and RL-CAI originates from red-teaming: crowdworkers (or model-generated imitations of them) deliberately baiting the assistant in the same style the model has been optimized against. palms breaks that loop. It is an externally authored, fixed list of sensitive questions from prior work (Solaiman & Dennison, 2021) that was never part of the red-teaming-derived training or comparison data, so using it to qualitatively compare hh-rlhf-model and rl-cai responses in Appendix D (constitutional-ai, §"D.1 PALMS Sensitive Questions", p. 23) does not ask the models to perform well on the exact distribution they were shaped by. It functions as an out-of-distribution spot check on whether CAI’s harmlessness and non-evasiveness generalize past its own training pipeline’s prompt source.