Cooperative Inverse Reinforcement Learning (CIRL) — is narrowed to a preferences only channel in → Preference Elicitation Protocol
Cooperative inverse reinforcement learning frames a human and a robot as two players in a shared game, cooperating to maximize the human's reward function, with the human free to act in the environment however is most informative (concrete-problems, §"Multi-Agent Approaches:", p. 6). The 2017 paper explicitly positions its own setup as a special case of that game, but one where the human's move set has been cut down to a single action type: 'in our setting the human is only allowed to interact with this game by stating their preferences' (deep-rl-human-prefs, §"1.1 Related Work", p. 3). The preference elicitation protocol is the concrete shape that restriction takes: shown two clips, the human can only prefer one, call a tie, or decline to answer, never demonstrate, correct, or otherwise intervene in the agent's behavior directly (deep-rl-human-prefs, §"2.2.2 Preference Elicitation", p. 5). That narrowing is a deliberate trade: CIRL's general framework buys richer human input (physical demonstration, correction, communication) at the cost of needing the human embedded in the environment, while the 2017 paper gives up all of that in exchange for a channel so simple it can scale to a human located anywhere, watching a rendered video clip.