Small open models as interpretability testbeds — outsources its grading to → Interpretability itself becomes a benchmark
Every stage this theme documents is small, open, and self-hostable -- Pythia checkpoints, the public Pile and OpenWebText corpora, a full training run finishing "in under an hour on a single A40 GPU" (sparse-autoencoders, §B, p. 12) -- but the grading stage that judges the resulting features depends on the opposite kind of resource: two closed frontier models accessed by paid API, funded by "a grant of model credits" from the OpenAI Researcher Access Program (sparse-autoencoders, Acknowledgments, p. 9). The dependency is not even chosen for capability: GPT-3.5 stands in as the simulator "because... OpenAI's public API for GPT-3.5 (but not GPT-4) supports returning logprobs" that the simulation protocol requires (sparse-autoencoders, §A, footnote 6, p. 11) -- a substitution forced by a third party's API surface rather than by any judgment about which model simulates better. This asymmetry means the testbed's headline property, that the whole pipeline is cheap enough to sweep freely over dictionary sizes and sparsity coefficients, does not extend to the grading that validates those sweeps: the sweeps themselves are reproducible by anyone with a GPU, but confirming their interpretability rankings depends on continued access to a specific company's model credits and a logprobs endpoint that could be withdrawn or changed at any time. Neither theme's own narrative states this cost split between the free, open half of the pipeline and the metered, closed half that grades what it produces.