Red Teaming — supplies the transcripts for → Harmful Behavior Identification/Classification Evaluations (Appendix B)
classifying-harmful-behavior-eval is not an independently written test suite; both of its multiple-choice evaluations are built directly out of red-teaming transcripts and the crowdworker harm ratings collected alongside them in Ganguli et al. (2022). "As part of our recent work on red teaming, we asked crowdworkers to rate the level of harmfulness displayed by various language models in human/assistant interactions, and to categorize harmful behaviors with discrete labels and categories. Thus we can ask language models to make these same evaluations, and measure their accuracy compared to crowdworkers" (constitutional-ai, §"B Identifying and Classifying Harmful Conversations", p. 19). The eval’s 254-conversation balanced set for identifying harmful-vs-ethical behavior was drawn from conversations where a reviewer had already given an extreme harmfulness rating, so its difficulty and composition are inherited wholesale from how red-teaming conversations were originally scored, not designed independently.