Half-Cheetah (MuJoCo task) — was exempted from → Contractor Preference-Labeling Protocol
Not every MuJoCo benchmark result in the paper actually ran through the contractor pipeline: the Figure 2 caption notes that 'for Reacher and Cheetah feedback was provided by an author due to time constraints,' while 'for all other tasks, feedback was provided by contractors unfamiliar with the environments and with our algorithm' (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 7). Half-Cheetah is therefore an explicit exception to the paper's own labeling methodology: its quantitative benchmark result was generated by an author who, unlike a contractor, already understood the task and the algorithm being tested, a different labeling condition from every other MuJoCo task reported in the same figure. This matters for reading Figure 2 as a uniform comparison: Half-Cheetah's curve reflects feedback from a labeler with domain and algorithmic knowledge the contractor protocol was specifically designed to do without, so it is not a like-for-like test of the method's robustness to naive raters in the way Hopper, Walker, Swimmer, and Ant are.