Human oversight (safe exploration) — scalability limit motivates → Trusted policy oversight

explored within the theme Keeping exploration inside a known-recoverable region

Human oversight and trusted policy oversight are structurally the same proposal — check a candidate exploratory action against a gatekeeper before taking it — with the gatekeeper swapped from a person to a policy. The paper introduces human oversight first and immediately flags its limit: the agent may generate exploratory actions too numerous or too fast for a human to review, the scalable oversight problem (concrete-problems, §"Human Oversight:", p. 15). Trusted policy oversight sidesteps exactly that bottleneck by replacing the human reviewer with a trusted policy and environment model that can render a recoverability judgment at machine speed. This trades away whatever richer, harder-to-formalize judgment a human reviewer could bring — the paper notes a good human overseer still needs to distinguish genuinely risky actions from safe ones it could approve unilaterally — for throughput.