Top-and-random vs. random-only interpretability scoring — is vindicated by → Interpretability-vs-kurtosis/skew correlation analysis

explored within the theme Interpretability itself becomes a benchmark

Appendix C.2 defends preferring top-random over random-only scoring with an empirical claim it does not yet prove there: 'the features we are considering are heavy-tailed, so with limited sample sizes, we should expect random samples to underestimate the true correlation.' Appendix C.4's kurtosis/skew analysis is where that claim actually gets tested rather than asserted, finding score correlates with skew (about 0.20) and kurtosis (about 0.15) but not with mean (-0.09) or variance (0.02), direct evidence that the dictionary's features really are heavy-tailed in the specific sense the sampling argument needed. Read in sequence, the case for trusting top-random scores over the stricter random-only alternative (Figure 9) is not just an intuition about rare features; by Appendix C.4, it is backed by a measured statistical property of the very features being scored. Without that later analysis, C.2's heavy-tailedness claim would be doing real argumentative work while remaining an unverified assumption.