bAbI tasks

Concrete Problems in AI Safetyinherited

A prior benchmark suite of prerequisite toy question-answering tasks (Weston et al. 2015), cited as a model for the kind of broad, diagnostic benchmark suite envisioned for safe exploration research.

Before large language models, Weston and colleagues built the bAbI tasks: roughly twenty tiny question-answering puzzles, each isolating one skill a reading system ought to have. Concrete Problems in AI Safety (2016) points to them as a model for the diagnostic benchmark suite that safe exploration research still lacks.

A section on the benchmark suite itself comes first, then one explaining that the paper borrows its structure for safe exploration research, not its technique.

A benchmark suite for prerequisite reasoning skills

The bAbI tasks are a benchmark suite introduced by Weston et al. (2015), consisting of around twenty synthetic question-answering tasks, each isolating one prerequisite skill for language understanding and reasoning: tracking a single supporting fact, combining two or three supporting facts, counting, negation, indirect reference, simple induction, and similar sub-problems. No single task is individually difficult; the point of the suite is that a model's coverage across all of them, rather than its score on any one, measures whether an architecture has learned generalizable reasoning rather than a task-specific shortcut. This concept predates Concrete Problems in AI Safety and is inherited by it as an example of good benchmark design (concrete-problems, §"6 Safe Exploration", p. 16).

A structural template, not a technique

Concrete Problems invokes bAbI only as a template for a benchmark Safe exploration research still lacks, not as a technique to adopt (see bAbI tasks — is the benchmarking template for → Safe exploration). The paper envisions an analogous suite of toy environments, covering conceptually distinct catastrophe types both physical and abstract, against which a safe-exploration architecture's coverage could be scored the way bAbI scores an architecture's coverage of reasoning sub-skills (concrete-problems, §"6 Safe Exploration", p. 16). The paper names the disanalogy itself: bAbI scores correctness of an output, while the envisioned suite would have to score the absence of a bad outcome, a harder target to benchmark because an agent can trivially avoid catastrophe simply by avoiding action altogether. This places bAbI in the corpus's broader pattern, documented in Turning an abstract oversight worry into dials and pools that can run together and The paper names exactly where its own fix stops working, of naming exactly where a borrowed idea stops fitting.