IC-162The majority of subject-verb agreement performance in Pythia-70m is explained by approximately 100 SAE feature nodes and in Gemma-2-2b by approximately 500 nodes, compared to approximately 1500 and 50000 neurons respectively

Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, Aaron Mueller

SourceSparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

The authors measure faithfulness and completeness of feature circuits versus neuron circuits for four subject-verb agreement structures. Small feature circuits (100 nodes for Pythia, 500 for Gemma) explain the majority of the model's agreement performance, while neuron circuits require roughly 15 times more nodes (1500 for Pythia) or 100 times more (50000 for Gemma) to explain only half the performance. Ablating just a few nodes from the feature circuits eliminates the model's task performance, even with all SAE error terms retained.

Evidence
correlational
Key metric
Pythia: ~100 feature nodes for majority of performance, ~1500 neurons for half; Gemma: ~500 feature nodes for majority of performance, ~50000 neurons for half
Caveat
SAE error nodes are high-dimensional and coarse-grained, so they cannot be fairly compared to neurons; the authors also report results with error nodes removed, showing that removing residual stream errors severely disrupts performance.
Model
Pythia Pythia-70m, Gemma 2 2B
Concepts
Explanation faithfulness
Related work
Bricken et al. 2023 [context], Cunningham et al. 2024 [context]
Related findings
IC-160, IC-161
Extraction
automatic-extraction