Scalable oversight — is addressed by adapting → Distant supervision
The paper proposes distant supervision as complementary to semi-supervised RL for scalable oversight, but flags it as the less experimentally ready of the two. Distant supervision is imported wholesale from semi-supervised and weakly-supervised learning in NLP, where the underlying assumption is that examples are i.i.d. — exactly the assumption an interactive RL agent violates. The text says explicitly that 'expanding these lines of work and finding a way to apply them to the case of agents, where feedback is more interactive and i.i.d. assumptions may be violated, could provide an approach to scalable oversight that is complementary to the approach embodied in semi-supervised RL' (concrete-problems, §"5 Scalable Oversight", p. 13). Unlike semi-supervised RL, which the paper scopes into concrete cartpole and Atari experiments, distant supervision is offered only as a direction still needing that adaptation.