Asimov's first law of robotics
Concrete Problems in AI Safety — inherited
A fictional robotics principle (robots must not harm humans) that Weld & Etzioni (1994) proposed as a target for formalization within classical AI planning.
A robot may not injure a human being. Asimov wrote his First Law as fiction in the 1940s, and half a century later Weld and Etzioni took it seriously enough to attempt a formalization within classical AI planning. This corpus remembers both moves.
The page tells that double story first, a fictional injunction and a serious attempt to formalize it, then explains why the corpus treats the law mainly as a foil.
A fictional injunction and a serious attempt to formalize it
Isaac Asimov's First Law of Robotics — "a robot may not injure a human being or, through inaction, allow a human being to come to harm" — appeared first as a plot device in his 1940s science fiction, not as an engineering specification. It predates the AI-safety research this corpus surveys by more than seventy years, and Concrete Problems cites it only through a serious academic attempt to take it literally: Weld and Etzioni's 1994 paper proposed formalizing the law within classical AI planning, the STRIPS-style framework in which an action is a set of preconditions and effects over predicate logic and a plan is a sequence of actions guaranteed to reach a goal state (concrete-problems, §"Other Calls for Work on Safety:", p. 21). Doing this seriously turns out to be hard for a specific technical reason: "harm" is not a primitive fluent that actions straightforwardly set or unset, and a planner that must avoid ever reaching any harmful state needs some tractable way to reason about the side effects of every action it considers, not just its intended effects.
Why the corpus treats it as a foil
Concrete Problems positions Asimov's law as the earliest entry in a twenty-year lineage of informal calls to formalize harm-avoidance, then immediately shows how far the field still had to go. The edge Asimov's first law of robotics — prefigures informally → Accidents in machine learning systems argues that the law names a goal without any mechanism for pursuing it: it is one monolithic injunction, where Accidents in machine learning systems fractures "harm" into five independently attackable technical causes (side effects, reward hacking, scalable oversight, safe exploration, distributional shift). A second edge, Formal verification (of cyber-physical systems) — supplies the rigor asimov lacked → Asimov's first law of robotics, pairs the law with Formal verification (of cyber-physical systems) of the federal aircraft collision-avoidance system as a picture of what rigorous, machine-checked harm-avoidance actually looks like — and then undercuts its own contrast, since the same section admits that this level of formal rigor "has not focused much on modern machine learning systems" (concrete-problems, §"Cyber-Physical Systems Community:", p. 20). Both edges sit in the A decades-old, informal idea resurfaces as an unsolved precursor connective theme, and the second also belongs to The paper names exactly where its own fix stops working: Asimov's aspiration, in other words, remains as unmet in 2016 as it was in 1994. The concept threads through the Framing safety as accidents, not superintelligence and Formal verification and control theory recur as safety building blocks themes and up into the Where Concrete Problems draws its own lines and Bounding what an agent can change or explore, not just what it's told to want superthemes, functioning everywhere as a marker of how much distance separates naming a safety goal from being able to build it.