Testing generative capability: translation and summarization benchmarks — resists mitigation alongside → The public NLP benchmark suite used to catch regressions

grounded in Training language models to follow instructions with human feedback · part of the supertheme The evaluation apparatus: measuring quality, capability loss, and safety together

Both groups catch the same alignment tax, but they do not recover uniformly once InstructGPT tries to fix it, and the mitigation experiment cuts across the two themes rather than respecting the boundary between them. Adding pretraining updates to PPO fine-tuning (PPO-ptx) reverses the regression on most of the public NLP suite and even pushes HellaSwag past plain GPT-3. It does not fix everything: after this mitigation, the PPO-ptx model still lags GPT-3 on DROP and SQuAD v2, from language-understanding-benchmarks' comprehension suite, and on the WMT 2015 French-to-English translation task from this theme's generation suite, with the paper conceding more work is needed on all three (instructgpt, §"4.2 Results on public NLP datasets", p. 15). The paper's own grouping of its unresolved cases is therefore by resistance to the fix, not by task type: a discrete-reasoning benchmark, an unanswerable-question benchmark, and a translation benchmark turn out to share one failure mode while the rest of both families recover.