Superposition and the case against the neuron basis — dictates the sparsity mechanics of → Unmixing activations with a sparsity penalty
Superposition's theory supplies more than motivation for building a sparse autoencoder at all -- it supplies the two numbers that fix the autoencoder's shape. The paper states plainly that "[f]eatures must be sufficiently sparsely activating for superposition to arise because, without high sparsity, interference between non-orthogonal features prevents any performance gain" (sparse-autoencoders, §"1 Introduction", p. 1); the L1 sparsity loss operationalizes that exact condition, and the expansion factor R, which sets dhid = R·din, operationalizes the companion claim that models pack more features than dimensions by giving the dictionary room to be genuinely overcomplete. The theoretical backing cited for the L1 term is a stronger promise than a design justification: Wright & Ma (2022) are cited for showing reconstruction with an L1 penalty "can recover the ground-truth features that generated the data" (sparse-autoencoders, §"2 Taking Features Out of Superposition...", p. 3), which implies a single correct decomposition exists to be recovered -- the same promise the sparsity-reconstruction sweep's smooth, knee-less curve across alpha then quietly sits in tension with. Neither theme states this outright: the diagnosis theme reports the recovery guarantee as settled background theory, and the instrument theme reports the knee-less sweep as a footnote about hyperparameter selection, but together they show the autoencoder's own hyperparameter sweep is quiet evidence against the very theoretical guarantee that licensed building it as an L1-penalized network in the first place.