The ladder of decomposition baselines — ranks every rung through → Interpretability itself becomes a benchmark

part of the supertheme Decomposing the model into features

Every rung of this ladder -- default basis, random directions, PCA, ICA, and the sparse dictionary itself -- is ranked by exactly one shared instrument, the GPT-4/GPT-3.5 autointerpretability score, so any blind spot in that instrument is a blind spot in the whole ladder at once rather than in any single method's score. The paper documents such a blind spot directly: because in the current protocol the scorer is "unable to verify outputs by looking at changes in output or other data," and because later features "are often best explained by their effect on the output" rather than by input text alone (sparse-autoencoders, §"3.2 Sparse Dictionary Features Are More Interpretable Than Baselines", p. 5), the scorer is structurally weaker at judging exactly the deep-layer, output-oriented features where interpretability differences between rungs might otherwise be sharpest. This is not a hypothetical risk: the sparse dictionary's own measured advantage over the ladder "declines as we move through the model, being comparable to ICA in layer 4 and showing minimal improvement in the final layer" (sparse-autoencoders, §3.2, p. 5) -- a pattern equally consistent with the dictionary's real advantage shrinking with depth or with the shared scorer simply losing power to detect it there. Because the whole ladder answers to one grader, the ranking's late-layer flattening cannot by itself distinguish those two explanations, a limitation neither the ladder theme, which treats the ranking as settled, nor the scoring theme, which treats its own controls as sufficient, states.