Scalable oversight — is operationalized as → Semi-supervised reinforcement learning
Concrete Problems introduces scalable oversight only informally, as the difficulty of ensuring safe behavior when a true objective is too expensive to evaluate on every training example. Semi-supervised reinforcement learning is the paper's attempt to turn that informal worry into a concrete, tunable quantity: an RL agent that observes ground-truth reward on only a small, controllable fraction of timesteps or episodes while still being evaluated on all of them (concrete-problems, §"5 Scalable Oversight", p. 11). Framing the problem this way converts a philosophical concern about oversight budgets into an experimental parameter — the labeled fraction — that can be dialed from fully supervised down to almost unsupervised and tested directly on toy domains like cartpole or Atari, rather than remaining a purely conceptual worry about what a designer 'should' be able to check.