Prior human-in-the-loop RL and why it didn't scale — fails the budget test set by → The economics of human feedback

part of the supertheme Deep RL from Human Preferences: the 2017 proof of concept

The economics theme's accounting gives a precise, quantitative reason why the prior approaches surveyed here stalled before reaching deep-RL scale, one the related-work discussion only gestures at qualitatively. Akrour et al.'s whole-trajectory preferences are the clearest case: the paper reports that its own segment-based method gathers "about two orders of magnitude more comparisons" than Akrour's approach, yet requires "less than one order of magnitude more human time" (deep-rl-human-prefs, §"1.1 Related Work", p. 3). Read against the economics theme's own metric, minutes of attention rather than number of queries, that gap says whole-trajectory preferences were never cheap per query; watching and judging a full trajectory costs far more than watching a one-to-two-second clip, so scaling up comparisons under that format would have scaled human time linearly rather than the sublinear way segments allowed. TAMER and the MacGlashan/Pilarski line fail the same budget on a different axis: because "learning only occurs during episodes where the human trainer provides feedback" (deep-rl-human-prefs, §"1.1 Related Work", p. 3), the human's attention is coupled to the agent's entire experience stream rather than to a bounded set of queries, economically indistinguishable from the per-timestep feedback the economics theme treats as too expensive to afford. None of the prior work is diagnosed this way in the paper's own text, but the economics theme's budget lens explains, in cost terms rather than capability terms, why each predecessor stayed confined to policies learnable "relatively quickly" or to a handful of hand-coded degrees of freedom (deep-rl-human-prefs, §"1.1 Related Work", p. 3).