Dictionary-feature patching procedure — is tested against → Indirect Object Identification (IOI) task
The Indirect Object Identification task is dictionary-feature-patching’s proving ground: a previously-studied, well-characterized model behavior (Wang et al., 2022) in which a model completes sentences like "Then, Alice and Bob went to the store. Alice gave a snack to ___" with the correct indirect object, chosen because it is well-understood externally even though the dictionaries used to explain it were trained with no knowledge of the task at all — "the training of our feature dictionaries does not emphasize any particular task" (sparse-autoencoders, §"4 Identifying causally-important dictionary features for indirect object identification", p. 6). Applying the patching procedure here, and comparing against an equivalent intervention using the residual stream’s PCA decomposition, shows dictionary features reach a given KL divergence from the target output using fewer patched features and a smaller mean edit magnitude than patching an equal number of PCA components (sparse-autoencoders, §"4.2 Precise localisation of IOI dictionary features", p. 6). That comparison is only informative because the feature subset patched is chosen by an automated ranking procedure rather than by hand, and IOI’s already-known circuit is what makes the ranking’s quality externally checkable at all: unlike the autointerpretability score, which has no independent ground truth to verify against, a task with a previously-published circuit lets the paper confirm that features patching identifies as causally important actually are.