Keeping exploration inside a known-recoverable region — is the only one of its mechanisms to persist into → Learning from demonstrations instead of a specified reward

hindsight · grounded in Training language models to follow instructions with human feedback · part of the supertheme Bounding what an agent can change or explore, not just what it's told to want

This theme proposes five mechanisms for physically constraining exploration: simulated exploration, bounded regions, trusted-policy oversight, human oversight, and use of demonstrations. Only the last has a documented afterlife in this corpus. It is exactly demonstration-based bounding, singled out in the source paper as reducing the need for risky exploration altogether, that the other theme traces forward into InstructGPT's supervised fine-tuning stage, trained directly on roughly 13,000 human demonstrations (instructgpt, §"3.1 High-level methodology", p. 6). Simulated exploration, bounded regions, and the two oversight mechanisms have no comparable descendant documented anywhere else in this wiki; they remain confined to the 2016 paper's own toolkit. Neither theme states this asymmetry on its own: this theme presents all five mechanisms as parallel options, and the demonstrations theme describes its genealogy without noting that it is the exception rather than the rule among this theme's proposals.