Random directions baseline — anchors → Autointerpretability score
To make the autointerpretability comparison fair to a metric built for nonnegative activations, the paper applies a specific preprocessing step only to the two baselines without a built-in nonlinearity: 'for the random directions and for the default basis in the residual stream, we replace negative activations with zeros so that all feature activations are nonnegative,' matching what a ReLU-based dictionary feature would produce natively. Random directions therefore are not scored on raw projections but on this clipped, half-space-restricted version, which if anything should make them easier, not harder, to interpret. That the random-directions floor still sits low despite this accommodation is what gives the autointerpretability score its usable dynamic range: every other decomposition, including the default basis, which Appendix C.2 finds is statistically no better than this floor in the residual stream, is judged against a baseline already given the benefit of the doubt.