Distant supervision — offers a structurally different route than → Hierarchical reinforcement learning (HRL)

explored within the theme Scaling oversight without a learned reward predictor

Distant supervision and hierarchical RL are presented as the two alternatives to semi-supervised RL for scalable oversight, and they weaken the oversight requirement along different axes. Distant supervision keeps the same fine-grained, per-decision timescale but accepts a weaker signal at each point — aggregate statistics or noisy rules instead of a verified per-example evaluation. Hierarchical RL keeps a strong, eventually-verifiable signal but changes the granularity at which it is required, since the top-level agent needs evaluation only on a small number of abstract, long-horizon actions while dense synthetic reward covers everything below it (concrete-problems, §"5 Scalable Oversight", p. 13). One branch makes each check cheaper; the other makes checks rarer — cheap-but-weak versus infrequent-but-strong, two orthogonal ways of stretching the same scarce oversight budget.