Impact regularizer (defined) — borrows its safe region from → Reachability analysis
Once the naive state-distance penalty is rejected for resisting all change, the paper needs a principled way to say what counts as a safe fallback the agent may deviate from. It reaches for reachability analysis, describing the refined, baseline-policy version of the regularizer as "somewhat reminiscent of reachability analysis... or robust policy improvement" (concrete-problems, §"Define an Impact Regularizer:", p. 5). Reachability analysis contributes the specific piece the naive version lacked: a formal notion of which states remain safely recoverable under a given policy, which is what lets "known safe (e.g. low side effect) but suboptimal policy" mean something precise rather than an ad hoc placeholder. The impact regularizer's baseline is only as rigorous as the reachability guarantee behind it; without that borrowed machinery, "safe but suboptimal" is just an assertion.