Connections the source papers never draw, visible only by reading them together
A large share of the connections in this store carry one label: links no source paper draws itself, visible only when one paper is read against an earlier or later one. Goodhart's law fragments into hedging and evasiveness without either paper naming it; demonstrations quietly become supervised finetuning once the reward-inference machinery drops away; and a later steering method's few-shot baseline reveals something neither paper states about what few-shot prompting is actually good for.
Most seams reach between an early proposal and the pipeline built years later, a screened sign-up sheet becoming a labeler-screening process, math surviving unchanged, an ad hoc clip server becoming a purpose-built interface, and an assumed error rate becoming a measured one. A further, smaller cluster reaches all the way to the corpus's newest paper, where an older framing about a model's baseline character gets a measurable sign and magnitude, and an older question about the cost of safety training finally gets an answer near zero.
Twenty-two of the twenty-three edges in the store are flagged as hindsight links: connections that exist only because a later paper can be read against an earlier one, not because either paper states the link itself, spanning from 2016 through 2022, into 2023 with CAA's addition, and now into 2025 with persona vectors'. Goodhart's law is empirically instantiated by reward overoptimization, and further specializes into excessive hedging and evasiveness, one 2016 principle fragmenting into several nameable 2022 phenomena; the edge between excessive hedging and evasiveness shows the fragments sharing one shape across a helpfulness/harmlessness axis, while the edge showing the KL penalty was not the fix applied to evasiveness demonstrates the two fragments needed different repairs. Use demonstrations is operationalized as supervised fine-tuning, shedding the reward-inference machinery the 2016 paper considered but keeping the core wager that demonstrations can substitute for a hand-specified objective. Scalable oversight is retroactively operationalized by the reward model, a match neither the 2016 paper nor InstructGPT itself draws, visible only once Constitutional AI names reward models as scaled supervision. And the Elo score generalizes win rate, one paper's single-anchor metric becoming the next paper's multi-snapshot instrument. A second cluster of seams runs from the 2017 paper straight to InstructGPT, five years apart, on matches neither paper draws itself. The Pong ablation's frozen, partially-trained reward predictor produces bizarre, exploit-driven behavior that anticipates reward overoptimization, InstructGPT's own reward model being gamed by a policy at a different scale and in a different domain, a resemblance InstructGPT never cites the earlier paper for. The 2017 paper's fallback for tasks with no reward function, ask a human whether generated behavior fulfills a natural-language goal, prefigures instruction following outright, the same wager that human judgment of goal-fulfillment can substitute for a reward nobody knows how to write down, five years before it became a primary training paradigm rather than a one-off grading method. The non-stationary reward problem the 2017 architecture was built around resurfaces under iteration in PPO training whenever InstructGPT's own RM-then-PPO loop is repeated rather than run once, reproducing the moving target the earlier paper's whole design responds to. And where the 2017 paper assumed a constant rater error rate as a modeling convenience, InstructGPT's inter-annotator agreement is what it looks like to go and measure that rate instead, one paper guessing what the other later checked. A third cluster reads the 2017 paper's machinery as ad hoc and InstructGPT's as its industrial replacement. Contractor preference labeling, an unscreened sign-up sheet with no quality check, becomes a screened operation in the labeler screening process, and the same 2017 protocol, an ad hoc clip server and keyboard interface, is formalized as InstructGPT's labeling interface, a purpose-built tool capturing structured judgments the 2017 setup only sketched. The Bradley-Terry model is carried over unchanged into the reward model loss, the identical 1952 pairwise-comparison mathematics still fitting a language-model reward five years after fitting a robot's; the asynchronous reward-learning architecture trades its concurrency for a fixed dataset in the reward model, InstructGPT training once on a static set of comparisons rather than refitting continuously against a moving policy. The trajectory segment generalizes into the RM dataset, the same design principle, judge the smallest unit a human can evaluate quickly, surviving the jump from a two-second clip to a full completion. Reward normalization prefigures the scale fix in reward model loss, since a comparison-only objective leaves scale undetermined in both papers, though 2022's fix anchors to an interpretable zero rather than an optimizer-convenient one; and the preference elicitation protocol avoids the overfitting problem solved by k-choose-2 batching only because its strictly pairwise design never had to face the correlated comparisons a K-wise ranking produces. A fourth seam runs the direction of time backward, from 2023 into 2022 rather than toward it: few-shot prompting shares its mechanism with the GPT-3-prompted baseline, but the two papers draw opposite lessons from using it, a resemblance visible only by comparing them side by side. InstructGPT treats a hand-optimized few-shot prefix, chosen by a 'prefix-finding competition' against its own reward model, as a serious lever worth pushing hard, since it was the only way to make raw GPT-3 look like it was following instructions at all. CAA finds the same technique comparatively weak a year later, losing out to a single system-prompt instruction for steering any of seven behaviors on a model already RLHF-tuned to follow instructions. Neither paper states the other's half of the comparison, but together they locate what few-shot prompting is actually good for: eliciting a response mode a model doesn't yet have, not fine-tuning a behavior a chat-tuned model can already exhibit. A fifth seam reaches from 2021/2022's HHH framing and InstructGPT's alignment tax all the way to 2025's persona vectors, on matches persona vectors itself never states as such. Evil trait inverts the baseline encoded by the HHH framework: the paper frames the Assistant persona as typically trained to be helpful, harmless, and honest before naming evil as its primary example of a trait models sometimes deviate toward, and Appendix J.4's CAFT comparison gives that framing a number, both base chat models tested exhibit strong negative projection values on the extracted evil direction, in contrast to hallucination's baseline sitting already near zero, so HHH's harmless criterion turns out to have a measurable sign and magnitude on the same activation-space axis the paper later uses to steer, monitor, and mitigate evil, a coordinate neither Askell et al. nor Bai et al. had the means to quantify. Persona vectors also answers a question InstructGPT posed and only partly solved: preventative steering answers alignment tax, the older paper's finding that RLHF's safety training measurably cost accuracy on public NLP benchmarks and needed pretraining-gradient mixing, PPO-ptx, to claw the loss back; preventative steering instead better preserves general capability than inference-time steering outright and, applied to a benign dataset that would not have shifted the persona anyway, has only a negligible effect on MMLU accuracy, a near-zero-tax outcome the 2022 paper never achieved without a separate patch. Underneath both hindsight seams sits one plainly-stated mechanism the lens now also carries as a member without the hindsight flag: supervised fine-tuning is the mechanism that produces finetuning shift, one epoch of rank-32 rsLoRA at a fixed learning rate on a single H100, a deliberately modest recipe held uniform across every trait-eliciting, emergent-misalignment-style, and benign dataset in the paper, so that differences in the resulting shift are attributable to what the data teaches rather than to how hard the model was trained. That edge is not a match between two papers a reader has to notice, persona vectors states it directly, but it is the mechanism the two hindsight seams beside it both depend on: the same low-rank finetuning is what evil trait's projection is measured against and what preventative steering must run during to answer InstructGPT's older tax question at all.