Two things built identically, pointed in opposite directions
A handful of mechanisms in this corpus turn out to be the same blueprint, built twice, and pointed at opposite ends. One measure gets minimized where its closest relative gets maximized; a single written constitution gets re-expressed in the opposite grammatical mood.
Security and abuse sit on opposite sides of one attack; an agent and the human teaching it swap which one becomes legible to the other; two InstructGPT failure modes wear the identical over-rewarded move as different disguises. Two robotics demonstrations later run that same trick on a single benchmark task, and the newest pair reaches across papers entirely: a sparse autoencoder and a steering vector turn out to be the same kind of direction in activation space, one reading it out unsupervised, the other writing it in by hand.
Several edges describe a pair that shares one underlying structure but runs it toward opposite ends. Security ML and abuse ML are mirror images defined by which side of an attack the system sits on; empowerment is minimized rather than maximized to build penalize-influence; cooperative inverse reinforcement learning and the reward autoencoder invert which party (agent or human) becomes legible to the other; unsupervised value iteration and unsupervised model learning apply an identical extraction move to RL's two dominant paradigms; SL-CAI's and RL-CAI's sixteen principles are the same constitution re-expressed in opposite grammatical moods, act-and-revise versus compare-and-choose. Two RLHF failure modes complete that set: excessive hedging and evasiveness are the same over-rewarded-proxy move wearing opposite surface disguises (more words versus fewer), and hallucination and excessive hedging bracket a single dial, how readily a model asserts what it cannot support, from opposite ends. The 2017 paper contributes two more pairs built the same way, same substrate, opposite setting. Hopper's ordinary benchmark and its backflip demonstration share an identical robot, dynamics, and observation space, but run that shared substrate to opposite ends of a specifiability axis: the standard task's reward is, by the paper's own description, a simple hand-writable quadratic, while the backflip demonstration extends the same physical setup past the point where any reward function is hand-engineerable at all, at barely more query cost than the benchmark it mirrors. The standard Enduro task and its keeping-pace demonstration share an identical game, action space, and learning algorithm; only the one sentence of instruction given to contractors changes, from rewarding overtaking to rewarding staying level with traffic, so the keeping-pace demonstration inverts the scoring objective of the Enduro task about as cleanly as an inversion can be staged, without touching a line of the game's code. A geometric mirror pair closes the set, drawn not within one paper but across two: sparse autoencoders and steering vectors are the same geometric object, a direction in activation space, built by opposite procedures and read in opposite directions. A sparse autoencoder's dictionary features are learned to read a direction's presence out of the residual stream, an unsupervised decomposition with no target concept specified in advance; a steering vector, whether CAA's Mean Difference or persona vectors' automated extraction, is built to write a chosen direction into the residual stream, always starting from a labelled concept and a set of contrastive examples. Persona vectors is the paper that states the pairing directly, describing sparse autoencoders as finding directions "in an unsupervised way" and naming that property a possible advantage over its own supervised pipeline. Reading out and writing in are the same underlying move, projecting onto or adding along a direction in the same vector space, run toward opposite ends: one recovers what is already there, the other inserts what wasn't.