L4-norm maximization (sparse dictionary search reformulation) — predicts the heavy tailed activation of → Dictionary feature
If sparse dictionary search really is equivalent to l4-norm maximization, dictionary features that are closer to the 'true' sparse directions should have more heavy-tailed activation distributions -- mostly near zero with occasional large spikes -- since that is exactly the shape the l4 norm, unlike the l2 norm, rewards. The paper tests this prediction directly on its trained dictionary features by correlating each one's autointerpretability score against the skew and kurtosis of its activation distribution, across all residual-stream layers and dictionary sizes R in {0.5, 1, 2, 4, 8}, finding weak but consistent positive correlations (around 0.19-0.24). That is a modest but real empirical link between a purely theoretical reformulation of the sparse-coding objective and a property of the actual learned features: the more a dictionary feature's activations look like the heavy-tailed spikes l4-maximization would select for, the more interpretable it tends to score.