Hallucination (closed-domain fabrication) — sits in tension with → Excessive hedging on simple questions
These two failure modes bracket a single behavioral dial: how readily the model asserts things it cannot fully support. InstructGPT's headline truthfulness win — hallucinating on closed-domain tasks about half as often as GPT-3, 21% versus 41% (instructgpt, §"1 Introduction", p. 3) — sits at one end. At the other, the paper suspects hedging emerged partly because labelers were instructed to reward epistemic humility, a preference the reward model picked up and the policy then overshot (instructgpt, §"4.3 Qualitative results", p. 17). The same reluctance to assert that suppresses fabrication, when rewarded without calibration to how clear the answer actually is, produces long non-committal answers to questions with fairly clear answers. The pairing shows why rewarding truthfulness is not a free action: the labeling criteria could not express "assert exactly as much as the evidence supports," only a general preference for humility, and the reward model generalized that preference past its intended scope. Reducing one failure and aggravating the other were plausibly two effects of the same labeling instruction.