A decades-old, informal idea resurfaces as an unsolved precursor

Robots should not harm humans, Asimov wrote in 1942, and the rule stayed unmechanized for decades, one injunction with no way for an engineer to act on it. Concrete Problems fractures that old rule into five causes an engineer can actually address.

From there the page turns to the 1969 frame problem, a knowledge-representation puzzle still unresolved when reinforcement learning inherited it, to formal verification's own claim to rigor where Asimov offered none, and closes on a 2011 thought experiment, the delusion box, that had already worked out wireheading's core failure.

Several of the corpus's concepts have an explicit ancestor from decades-earlier scholarship, whether an informal principle that only foreshadowed the problem or an already-formal model borrowed wholesale, long before a modern paper gave the problem its own technical treatment. Asimov's first law prefigures accidents in ML informally, a single monolithic injunction with no mechanism for an engineer to act on, which the modern paper immediately fractures into five attackable causes. The frame problem supplies the classical diagnosis for negative side effects, a 1969 knowledge-representation difficulty resurfacing, largely unaddressed, in reinforcement learning objective design. Formal verification supplies the rigor Asimov's first law lacked, showing what actually rigorous harm-avoidance looks like once a community commits to it, for cyber-physical systems if not yet for modern ML. And the delusion box gives a minimal formal model of environmental embedding, a 2011 result that predates Concrete Problems and strips the wireheading failure down to a stark, already-demonstrated case. A different kind of ancestor supplies the 2017 paper's reward-learning loss outright rather than merely anticipating it: the Bradley-Terry model, a 1952 statistical model for estimating score functions from pairwise comparisons, is adopted unmodified to supply the loss for the reward model, with no engineering needed beyond swapping in trajectory segments for whatever the original model ranked. The paper reaches for a second decades-old idea just to explain the first: the Elo score, chess's rating system, supplies the explanatory analogy for the Bradley-Terry model, licensing the reward network's output to be read as an interval-scaled rating rather than a bare classifier score, even though the paper never computes anything resembling an actual Elo rating.