The score is itself audited

A metric that grades itself as trustworthy is not the same as a metric that has actually been checked, and this paper spends real effort making sure its central score is the second kind rather than the first.

These nine connections trace how the autointerpretability score gets pressure-tested from several directions at once: a floor set by random directions, two separate fairness controls against its strongest rivals, and an explanation of what the score actually rewards. One honest admission stays attached throughout: the score can't see a feature's causal effect on the model's output, so it remains an imperfect stand-in even after all that scrutiny.

Nine edges show the paper interrogating its own metric rather than trusting it outright. The random-directions baseline anchors the autointerpretability score's floor, given every accommodation (zeroed negative activations) so that floor is not artificially depressed. Two fairness controls test separate confounds: the top-K baseline control density matches ICA to rule out a geometric asymmetry, distinct from the confound top-random scoring's own fragment-sampling compromise raises, which the paper checks with a random-only rerun. The kurtosis-skew correlation analysis explains what the score actually rewards, heavy-tailed activation rather than raw magnitude, supplies empirical evidence for sparse dictionary learning's L4-norm theory, and in turn vindicates top-random scoring's heavy-tailedness assumption. L4-norm maximization predicts this same heavy-tailed signature directly in dictionary features. Yet monosemanticity itself is only imperfectly proxied by the score, which cannot see causal effects on outputs — an imperfect instrument the paper uses anyway, with its blind spots stated plainly.