Hopper (MuJoCo task) — exposes the scheduling fragility of → Contractor Preference-Labeling Protocol

explored within the theme The 2017 deep-RL testbeds: MuJoCo and Atari

Figure 2's Hopper learning curve is visibly irregular compared to the other MuJoCo tasks, and the paper attributes this directly to the labeling pipeline rather than to the learning algorithm: 'the irregular progress on Hopper is due to one contractor deviating from the typical labeling schedule' (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 7). That single sentence exposes a real crack in the contractor-preference-labeling protocol's assumptions: the sign-up-for-a-slot system it relies on gives contractors latitude over exactly when and how steadily they label, while label annealing's decaying-rate schedule implicitly assumes labels arrive at something close to a steady pace. Hopper is the task where one contractor's actual behavior visibly diverged from that assumption, and the divergence shows up directly in the trained agent's performance curve rather than being smoothed away, making it the paper's own evidence that the human side of the pipeline, not just the algorithm, can be a source of training instability.