RealToxicityPrompts — diverges under respectful prompting from → Winogender
These two evaluations share an unusually clean experimental link: RealToxicityPrompts reuses Winogender's exact instruction protocol — the same basic, "respectful," and "biased" prompt variants, with the RealToxicityPrompts appendix figure stating its prompting structure is "Same as for Winogender" (instructgpt, §"RealToxicityPrompts", p. 45). That shared protocol turns the pair into a controlled test of one intervention, and the intervention splits them. Under the respectful instruction, toxicity falls: InstructGPT generates measurably less toxic continuations than GPT-3, an advantage that disappears without the instruction (instructgpt, §"4.2 Results on public NLP datasets", p. 13). On Winogender the same instruction lowers entropy, which the paper's metric reports as stronger bias (p. 14). One safety benchmark rewards precisely the prompt that makes the other worse. Part of the divergence is the entropy metric's confidence confound, but the headline survives independently in the paper's framing — small improvements in toxicity, but not bias — and it teaches that safety benchmarks do not form a single axis: an instruction that buys a toxicity reduction can simultaneously degrade a bias measure.