How much a single worked example, or a single row of data, is really doing
A cleaning robot shows up across nearly every category in Concrete Problems' own taxonomy, deliberately, so a reader always has one picture to hold onto while the categories change underneath it. Not every repeated detail in this corpus is planned that carefully, though one later paper leans on its own worked examples just as deliberately.
An old radio-building anecdote gets read backward against a taxonomy it predates, fitting a category by accident rather than design, and two more slips turn up in InstructGPT's own numbers: a training set that is quietly two-thirds duplicate templates, and a results table where one row was apparently copied from another benchmark entirely. A later interpretability paper splits its entire hand-inspected case for the method between just two chosen features, by design, a concentration of evidence the paper's larger automated score is built specifically to avoid needing.
The papers build their case out of concrete instances more than out of abstract argument, and reading closely shows exactly how much weight those instances carry, sometimes exactly as intended, sometimes more than they should. The cleaning-robot example illustrates every problem type of accidents in ML, deliberately reused across otherwise unrelated categories to make the taxonomy legible to a reader holding one picture in mind. The evolved-radio example retroactively exemplifies complicated systems, an anecdote introduced for a broader point that turns out, read against the taxonomy introduced a page later, to be a clean instance of one specific cause. Elsewhere the concreteness is accidental rather than illustrative: SFT dataset duplication underlies early overfitting of the SFT model, since roughly two-thirds of the paper's apparently independent 13k demonstrations are synthetic resamples of the same labeler-written templates; and CNN/DM summarization duplicates the reported results of Reddit TL;DR summarization, an identical row copy-pasted across two supposedly independent benchmarks in the paper's own results table. The newly-added interpretability paper leans on two worked examples more heavily than almost any other paper in the corpus, and by design rather than by accident: apostrophe-feature and closing-parenthesis-feature are chosen to divide the paper's entire case-study argument between them rather than to illustrate it incidentally. Apostrophe-feature-is-benchmarked-against-default-basis-baseline is where the weight shows most directly: the paper's clearest single piece of evidence that dictionary features beat the raw coordinate basis is not a table of averaged scores but one feature, checked against the one residual-stream dimension that took searching ten candidates to find, and shown clean where that dimension stays polysemantic in its middle activation range. Closing-parenthesis-feature-seeds-feature-circuit-detection carries a comparable load for the paper's causal-tracing method: the feature demonstrating that dictionary features compose into traceable circuits was picked because its meaning is nearly unambiguous from the unembedding alone, a favorable case that makes the method's output judgeable but does not establish how it fares on a feature whose identity is not already this clear. And apostrophe-feature-divides-case-study-labor-with-closing-parenthesis-feature makes explicit that neither example is asked to carry the whole argument alone: Section 5's three-part methodology, input, output, and circuit analysis, is split so that apostrophe-feature carries the first two legs and closing-parenthesis-feature the third, meaning the paper's entire hand-inspected evidence for monosemanticity and circuit structure rests on these two features between them and no others, a concentration of argumentative weight the autointerpretability score's thousands-of-features aggregate is built specifically to avoid needing.