Uncertainty-Based Query Selection (Ensemble Variance) — is a crude approximation of → Expected-Value-of-Information Query Selection

explored within the theme The economics of human feedback

The paper's authors are explicit that ensemble-variance querying is 'a crude approximation' to the query-selection problem, not a considered solution to it: the criterion they call more principled is expected value of information, credited to Akrour et al. (2012) and Krueger et al. (2016), which they leave entirely to future work rather than implement (deep-rl-human-prefs, §"2.2.4 Selecting Queries", p. 6). The gap between the two is not merely theoretical: the same passage adds that the ablation experiments show the ensemble-variance heuristic 'in some tasks actually impairs performance,' meaning the crude approximation can do worse than the alternative baseline the ablations actually test, uniform random querying. This matters because uncertainty-based-query-selection is the method every economics-of-feedback calculation in the paper (700 labels for MuJoCo, 5,500 for Atari) is built on top of, while the criterion its own authors regard as correct was never implemented, let alone benchmarked against it directly. The comparison leaves the paper's query-selection choice openly provisional rather than settled, a rare case where the authors flag their own method as possibly counterproductive before any reviewer does.