Keeping exploration inside a known-recoverable region — reaches the same safety envelope by a different route than → Penalizing side effects by distance from a baseline
Bounded exploration and impact regularization are built for different failure modes, exploration risk versus side effects, yet both cash out as the same control-theoretic move borrowed from prior work: bound the worst case rather than model consequences exactly. Impact regularizers measure distance from a baseline policy built via reachability analysis and robust policy improvement (concrete-problems, §"Define an Impact Regularizer:", p. 5); bounded exploration checks whether a candidate action would exit a safe region, a check related to H-infinity control's minimization of worst-case response to bounded disturbance (concrete-problems, §"Bounded Exploration:", p. 15). Neither theme's own account states the parallel directly: impact regularization treats reachability analysis as a specification aid for penalizing change already made, while bounded exploration treats H-infinity control as a fence against action not yet taken. The convergence is structural, not causal, two independently motivated bounding strategies reaching the same robust-control guarantee from opposite temporal directions.