From an open problem to reward models to AI feedback — is only half operationalized by → InstructGPT's three-step training pipeline

part of the supertheme Scalable oversight made real: from 2016 proposals to InstructGPT's RLHF pipeline

The lineage theme's own arc runs from an open problem to a human reward model to AI feedback, but the three-step pipeline theme's member concepts, GPT-3, SFT model, supervised fine-tuning, reward model, reward-model loss, RLHF, PPO training, stop at the first of those two implementations. Nothing among them varies who or what supplies the preference label; the label is human throughout every step. RLAIF, the feedback model, the hybrid human/AI preference model, and iterated online training, the lineage theme's later concepts, describe a second implementation, Constitutional AI's, that sits entirely outside the three-step pipeline's own concept set. The pipeline theme is therefore not the lineage fully built but its first chapter only; the lineage's claim that scalable oversight gets scaled again through AI feedback describes a different, sibling pipeline the three-step theme does not itself contain.