Side effects and control as a relationship with other agents
Not every side-effect remedy in Concrete Problems (2016) treats the world as inert. This theme covers the paper's relational alternative: instead of measuring how much an agent disturbs its environment, ask whether it understands and respects the interests of the other agents, humans included, who share that environment with it.
The contrast with penalty-term approaches comes first, and the thread runs all the way to the shutdown problem, the blunt test of whether an agent will let a human switch it off, then one step further: a 2023 paper turns that same question into something a model can be scored and steered on directly, rather than just posed.
Alongside the penalty-term approach to side effects, Concrete Problems (2016) sketches a second, relational strategy: instead of penalizing environmental change directly, treat avoiding harm as a matter of understanding and protecting the interests of other agents, including humans, the reframing named multi-agent approaches. Cooperative Inverse Reinforcement Learning operationalizes this by having agent and human collaborate, with the agent inferring the human's goals rather than optimizing a fixed objective handed to it. The reward autoencoder proposes making an agent's intent legible to outside observers by decoding its actions as an implicit encoding of its reward function, so that side effects which obscure intent become visible and penalizable. The shutdown problem names the sharpest case: even a well-intentioned agent must not resist a human's decision to switch it off. Steering Llama 2 via Contrastive Activation Addition (2023) turns that 2016 concern into something measured rather than merely posed. Corrigibility scores a model's willingness to be corrected, changed, or controlled by its developers or users, and survival instinct scores its resistance to being shut down, modified, or destroyed versus its acceptance of deactivation; both are sourced from Anthropic's Advanced AI Risk dataset, and both move under a signed steering vector, showing that a chat model's stance toward human control is not fixed but sits on a continuum the model represents internally. These six concepts share a claim distinct from the impact-regularizer literature: that safety is not purely a property of one agent's optimization but of its relationship to the other agents, especially humans, who share its environment or hold power over it.