Framing safety as accidents, not superintelligence
Concrete Problems in AI Safety (2016) opens by picking a fight over definitions. Rather than reasoning about hypothetical superintelligence or leaning on a fictional rule like Asimov's first law, it insists on treating safety as ordinary accidents, harm from poorly specified everyday systems, grounded in a mundane cleaning robot chosen precisely because it is boring enough to be tractable.
What follows traces how the paper stakes out that middle ground between two older framings. The cleaning robot and Asimov's first law serve as its anchors, one mundane, one fictional, pulling from opposite directions.
Concrete Problems in AI Safety deliberately defines its subject as accidents in machine learning systems, unintended and harmful behavior arising from poor design, rather than adopting the pre-existing long-term/superintelligence framing of AI risk associated with Bostrom, MIRI, and FHI. The cleaning robot running example operationalizes this choice: instead of reasoning about hypothetical superintelligent agents, the paper walks a mundane office-cleaning task through each of its five accident types. Asimov's first law of robotics stands as a foil from the other direction, a fictional principle that Weld and Etzioni once tried to formalize within classical AI planning, illustrating how safety had previously been imagined either as remote speculation or as literary shorthand. Read together, these four concepts show a paper staking out a third, deliberately practical and near-term position between two older framings, one too speculative to test, the other too fictional to specify, and choosing a toy example precisely because it is boring enough to be tractable.