Scaling oversight without a learned reward predictor — prefigures the rule driven labeling in → From an open problem to reward models to AI feedback
This theme's two branches fare unevenly against the lineage's later step. Hierarchical decomposition, a top-level agent delegating abstract goals to sub-agents pursuing denser synthetic rewards, has no realized analogue anywhere in the lineage: RLHF and RLAIF remain architecturally flat, a single policy against a single preference model, with no goal-delegating hierarchy. Distant supervision fares differently: its core move, substituting a noisy labeling rule for per-example evaluation, exemplified by DeepDive asking users to supply rules that generate weak labels (concrete-problems, §5 Scalable Oversight, p. 13), is what the feedback model in Constitutional AI's RL from AI Feedback actually does. The feedback model is prompted with a single constitutional principle alongside each response pair and produces a preference label from that principle rather than from independent per-example human judgment (constitutional-ai, §4.1 Method, p. 10), making it a rule-driven labeler in the distant-supervision sense, not a scaled-up version of per-example oversight.