CNN/Daily Mail summarization — duplicates the reported results of → Reddit TL;DR summarization
The two summarization benchmarks were meant to be independent checks — news articles with professional reference summaries versus Reddit posts with author-written ones — but the paper's reporting collapses them into one. In Table 14, the CNN/DM and TLDR rows are numerically identical across all twelve models, to three decimal places (0.182 through 0.220), which is implausible for two distinct datasets and almost certainly a duplicated row (instructgpt, §"E.1 Performance on public NLP datasets", p. 56). The copy-paste trail extends into Appendix D: the dataset-features boxes for both benchmarks describe their few-shot setting as containing "15 additional French / English pairs" — text carried over from the WMT translation figure (instructgpt, §"CNN/DM Summarization", p. 49; §"TLDR Summarization", p. 50). The practical lesson for a reader: the paper's ROUGE-L evidence on summarization is one measurement, not two independent confirmations, and at least one of the two reported rows cannot be what its label claims. Anyone citing InstructGPT's summarization scores should first establish which dataset the surviving numbers actually belong to.