How much a single worked example, or a single row of data, is really doing
A cleaning robot shows up across nearly every category in Concrete Problems' own taxonomy, deliberately, so a reader always has one picture to hold onto while the categories change underneath it. Other repeated details in this corpus are not so carefully planned.
An old radio-building anecdote gets read backward against a taxonomy it predates, fitting a category by accident rather than design, and two more slips turn up in InstructGPT's own numbers: a training set that is quietly two-thirds duplicate templates, and a results table where one row was apparently copied from another benchmark entirely.
The papers build their case out of concrete instances more than out of abstract argument, and reading closely shows exactly how much weight those instances carry, sometimes exactly as intended, sometimes more than they should. The cleaning-robot example illustrates every problem type of accidents in ML, deliberately reused across otherwise unrelated categories to make the taxonomy legible to a reader holding one picture in mind. The evolved-radio example retroactively exemplifies complicated systems, an anecdote introduced for a broader point that turns out, read against the taxonomy introduced a page later, to be a clean instance of one specific cause. Elsewhere the concreteness is accidental rather than illustrative: SFT dataset duplication underlies early overfitting of the SFT model, since roughly two-thirds of the paper's apparently independent 13k demonstrations are synthetic resamples of the same labeler-written templates; and CNN/DM summarization duplicates the reported results of Reddit TL;DR summarization, an identical row copy-pasted across two supposedly independent benchmarks in the paper's own results table.