SFT dataset — bootstrapped the existence of → API prompt distribution

explored within the theme Building the human side of the RLHF pipeline

The API prompt distribution looks like exogenous real-world demand, but it is partly a product of the dataset it now feeds. The Playground traffic the paper harvests was submitted to an earlier InstructGPT model trained via supervised learning on a subset of the demonstration data (instructgpt, §"3.2 Dataset", p. 6) — and those demonstrations exist because labelers were asked to write plain, few-shot, and use-case-based prompts themselves, precisely because instruction-style prompts were not often submitted to the regular GPT-3 models on the API (instructgpt, §"3.2 Dataset", p. 7). The circularity matters for interpreting the paper: users learned to write instruction-style prompts by interacting with a model that had already been taught, by labelers, what instructions look like. The distribution InstructGPT is trained and evaluated on is therefore not independent of the intervention; it is a co-evolved artifact — demonstrations shaped a model, the model shaped user behavior, and user behavior became the training and evaluation distribution.