Interpretability-vs-kurtosis/skew correlation analysis — supplies empirical evidence for → Sparse dictionary learning / sparse coding
Qu et al. (2019) proved, as pure theory, that searching for sparse overcomplete dictionaries can be reformulated as searching for directions that maximize the L4 norm, a measure closely related to kurtosis, the fourth statistical moment. The kurtosis/skew correlation analysis is the paper's empirical check on whether that theoretical reformulation shows up in practice: if sparse dictionary learning is, mathematically, a search for heavy-tailed directions, then real dictionary features recovered by an L1-penalized autoencoder should be more interpretable precisely to the degree that they are more heavy-tailed, which is what Table 2 finds, since skew and kurtosis correlate with score while mean and variance do not. In other words, this correlation analysis is the closest thing in the paper to a direct empirical fingerprint of sparse coding actually working as the L4-norm theory predicts, rather than an incidental property of the learned features. It also supplies the paper's account of why ICA, a technique from outside the sparse-dictionary-learning lineage that nonetheless explicitly maximizes non-Gaussianity, ends up as by far the strongest non-SAE baseline: it is optimizing for the same heavy-tailedness this analysis shows the SAE's dictionary features share.