Sparsity-reconstruction tradeoff (no single correct decomposition) — rules out a single canonical dictionary for → Sparse autoencoders (SAEs)

explored within the theme Unmixing activations with a sparsity penalty

Because sweeping alpha produces a smooth curve with no 'knee' or distinguished bend (Appendix B, Figures 6-7), the paper reads this as weak evidence against there being one uniquely correct sparse decomposition of a given activation space -- every (alpha, R) setting on the curve is a defensible, merely different, decomposition rather than an approximation to some single ground-truth dictionary. That is part of why the method is presented and evaluated across a range of alpha and R values rather than as a search for the 'right' hyperparameters: Sections 3 and 4 report results at multiple settings (e.g. alpha=.00086, R=2 for the main interpretability comparison, and a separate alpha sweep for the IOI patching results) instead of tuning toward one canonical dictionary. It also reframes what 'the model's features' means for sparse autoencoders as a method: not a fixed list to be recovered exactly, but a family of similarly-valid sparse bases that trade sparsity and fidelity against each other.