Ant (MuJoCo task) — hard codes an upright priority into → Contractor Preference-Labeling Protocol
The Ant task's reward-shaping result is often summarized as humans 'implicitly' preferring upright posture, but the contractor instructions actually make the priority explicit rather than implicit: raters were told 'the first priority is for the robot to be standing upright, and failing that for the center of the robot to be as high up as possible,' with forward speed demoted to a tiebreaker used only when both robots are equally upright or equally fallen (deep-rl-human-prefs, §"B.1 MuJoCo", p. 16). That ordering was a deliberate editorial choice built into the contractor-preference-labeling brief itself, distinct from the task's true reward function, which encoded uprightness only as a smaller additive bonus rather than as a lexicographic first priority (deep-rl-human-prefs, §"3.1.1 Simulated Robotics", p. 7). So the shaping advantage human feedback produced on Ant did not emerge from raters independently discovering that posture mattered; it was written directly into the one-to-two-sentence instructions they were given, making the labeling protocol itself, not just rater judgment, the source of the better-shaped signal.