Side effects and control as a relationship with other agents — passes on only its demonstration half to → Learning from demonstrations instead of a specified reward

hindsight · grounded in Training language models to follow instructions with human feedback · part of the supertheme Bounding what an agent can change or explore, not just what it's told to want

CIRL belongs to both themes because it does two things at once: it is a demonstration-adjacent technique for avoiding hand-specified rewards, and it is a relational framing in which safety depends on modeling another agent's goals and interests, a framing this theme also applies to the shutdown problem and to the reward autoencoder's legibility proposal. InstructGPT's supervised fine-tuning stage, the concrete descendant the other theme traces from demonstration-learning, inherits only the first half. It clones behavior from roughly 13,000 static human demonstrations offline (instructgpt, §"3.1 High-level methodology", p. 6); it does not query a human interactively, infer a latent goal, or carry any analogue of the shutdown problem or reward-autoencoder legibility that CIRL's relational half was meant to address. The genealogy from demonstration-learning to SFT is real but partial: the collaborative, human-modeling character this theme built around CIRL has no realized counterpart in the pipeline the other theme documents.