Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Sparse Autoencoders Find Highly Interpretable Features in Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-808
Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layers
IC-809
Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergence
IC-810
Individual sparse autoencoder features in Pythia-70m-deduped are monosemantic and have predictable causal effects on output logits, as demonstrated by an apostrophe feature whose ablation primarily suppresses the 's' token