Feedback Model — supplies the harmlessness half of → Hybrid Human/AI Preference Model

explored within the theme From an open problem to reward models to AI feedback

The hybrid preference model is trained on two datasets of different provenance stitched into one comparison set: 135,296 human-labeled helpfulness comparisons and 182,831 harmlessness comparisons generated entirely by the feedback model, one per SL-CAI prompt (constitutional-ai, §"4.2 Datasets and Training", p. 11). The mixing happens purely at the level of training data — the resulting model is architecturally an ordinary reward model, identical in recipe to the one used for pure human-feedback RLHF — so the innovation is confined entirely to where half of the comparisons came from, not to any change in how the model is built or trained. Once trained, the hybrid PM cannot itself distinguish which of its two source datasets a given judgment traces back to; helpfulness and harmlessness are fused into a single scalar reward.