bAbI tasks — is the benchmarking template for → Safe exploration
bAbI is invoked not as a technique but as a structural template for a benchmark that safe exploration research still lacks. bAbI tests language understanding by decomposing it into a broad set of small, diagnostic tasks, each isolating one prerequisite skill (counting, negation, indirect reference, and so on), so that a single architecture's coverage across the whole suite becomes a meaningful measure of progress (concrete-problems, §"6 Safe Exploration", p. 16). The paper envisions an analogous suite of toy environments covering conceptually distinct catastrophe types, physical and abstract, against which a safe-exploration architecture could be scored the same way. The disanalogy is instructive too: bAbI scores correctness of an output, while the envisioned suite would have to score the absence of a bad outcome, a harder thing to benchmark since an agent can trivially avoid catastrophe by avoiding action altogether.