Scalable oversight — is retroactively operationalized by → Reward model (RM)
Concrete Problems names scalable oversight in 2016 but supplies no worked instance at the scale of a deployed system; InstructGPT's reward model, trained in 2022, is never described by InstructGPT itself as an answer to that problem, since InstructGPT never uses the term. The connection only becomes visible once Constitutional AI revives the same idea under the label 'Scaling Supervision' and states that 'work on reinforcement learning from human feedback...has already taken a step in the direction of scaled supervision, since the reward signal in RL actually comes from an AI preference model (PM) rather than from immediate human oversight' (constitutional-ai, §"Scaling Supervision", p. 2). Read backward through that later framing, the reward model is exactly the cheap, densely-queryable proxy for an expensive true objective that scalable oversight called for six years earlier — a retrospective match neither the 2016 paper nor InstructGPT itself draws explicitly.