Decomposing the model into features

Eleven themes turn out to be one argument told in stages: first a diagnosis of why a model's internals resist being read directly, then a tool built in response, then ways of checking whether that tool earns its keep.

This supertheme follows that argument from the case against reading neurons directly, through the sparse autoencoder built to replace them and the older ideas it descends from, to the scoring and baseline comparisons that test it, the interventions that prove its features are causally real, and the safety ambition the whole project is offered in service of, sharing its method for reading model internals with the corpus's steering research.

This supertheme reads eleven themes from Cunningham et al. (2023) as one program moving from diagnosis to instrument to measurement to causal proof to safety ambition, sharing internals-reading machinery with the corpus's steering papers. Superposition and the case against the neuron basis supplies the diagnosis: models pack more features than they have dimensions, so neurons and default-basis coordinates are the wrong unit of analysis. Unmixing activations with a sparsity penalty supplies the instrument, a small autoencoder trading sparsity against reconstruction; the dictionary-learning lineage traces its descent from Yun et al. and Sharkey et al.; open models as interpretability testbeds names the small, open Pythia/Pile/OpenWebText infrastructure it is tested on. Automated interpretability scoring and the ladder of decomposition baselines supply the measurement, judging features against PCA, ICA, random directions, and the default basis. Causal proof by patching, and its companion single features under the microscope, supply the causal proof, editing features and reading two case-study features from both ends. Feature enumeration as a safety audit supplies the ambition: the resulting catalogue as a route to model audit. Reading the residual stream and steering as audit, the last two members, share that internals-reading machinery, also belonging to the steering-and-model-internals supertheme.