Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Scaling and evaluating sparse autoencoders
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-525
GPT-2 small's residual stream at layer 8 decomposes into two sub-spaces of approximately 25% and 75% of the dimensionality
IC-526
GPT-2 small's first token position has residual stream norms more than an order of magnitude larger than all other positions
IC-527
A 16 million latent sparse autoencoder substituted into GPT-4 yields a language modeling loss corresponding to 10% of GPT-4's pretraining compute