SQuAD v2 — forms the regression testbed with → DROP (Discrete Reasoning Over the content of Paragraphs)
Of the eight understanding benchmarks in InstructGPT's suite, these two were quietly promoted from evaluation to tuning: both the pretraining-coefficient sweep and the KL-coefficient sweep were judged on "SQuADv2 and DROP (the datasets we used for testing)" (instructgpt, §"4.2 Results on public NLP datasets", p. 15). The pair behaves as a single instrument throughout. A pretraining loss coefficient of 20 or more recovers their regressions on the 1.3B model, and the adopted value of 27.8 was validated against them (instructgpt, §"E.6 Fixing regressions on public NLP datasets", p. 53); raising the KL coefficient even 100-fold never fully recovers either (p. 15); and doubling RL training to 512k episodes re-opens the regression on exactly these two, with performance starting above GPT-3 and sinking back below it (instructgpt, §"E.6 Fixing regressions on public NLP datasets", p. 56). The consequence is easy to miss: "the alignment tax was mostly mitigated" operationally means "F1 was restored on these two extractive QA sets," so the claim inherits whatever the pair idiosyncratically demands — exact-span answers scored by token overlap — rather than measuring general capability loss.