Clip-Length Effects on Human Evaluation — predicts but was never tested on → Qbert (Atari game)
The clip-length-effects finding, that short clips take a rater disproportionately long to interpret relative to the information they carry, was tested empirically only on the MuJoCo robotics tasks, since the Atari reward model conditions on a stacked window of consecutive frames and so has no length-1 ablation to run there (deep-rl-human-prefs, §"3.3 Ablation Studies", p. 10). Qbert is the one place in the Atari results where the same underlying difficulty is invoked without ever being confirmed by an ablation: the paper attributes its outright failure to learn from real human feedback to the possibility that 'short clips in Qbert can be confusing and difficult to evaluate' (deep-rl-human-prefs, §"3.1.2 Atari", p. 8). The hedge, 'this may be because,' is the tell: the authors extend a mechanism they verified only on continuous-control tasks to explain a discrete-domain failure they never ran the corresponding experiment to check, making Qbert the paper's single suspected, but formally untested, instance of clip length actually costing the method a benchmark.