Failed routes left visible

Not every choice in this paper worked on the first try, and rather than quietly editing those attempts out, the paper leaves the dead ends standing next to the results that succeeded.

These seven connections gather that honesty: an architecture that had to change once it met a harder setting, an exception that turned out not to rescue the case it was raised for, a case-study feature that doesn't quite live up to its own billing, a scoring advantage that thins out with depth, and an early causal method abandoned once a better one replaced it, each limitation stated plainly enough to make the surrounding successes legible.

Seven edges show the paper's negative results doing real work. MLP SAE training stress tests sparse autoencoders and breaks two working assumptions at once, forcing tied-weights SAE to give way to an untied variant in MLP SAE training just to recover some of the lost capacity. Outlier dimensions fails to elevate the default-basis baseline above the random-directions floor, closing off an expectation the paper itself raised. The apostrophe feature demonstrates the limits of monosemanticity, since it does not cover every apostrophe, and dictionary feature scores unevenly across depth in the autointerpretability score, its advantage thinning toward later layers. A weight-based connection attempt is the failed precursor to feature circuit detection, its cosine-similarity shortcut abandoned once found to measure nothing the network's own computation respects. And sparse autoencoders runs counter to the depth profile of residual stream once a later paper's behavioral-separability curve is set against it. None of these limits are hidden; each is stated plainly enough to make the positive results beside them legible.