Light Dark Sparse autoencoder / Sparse autoencoders / K-sparse autoencoder / Topk sparse autoencoder / Scaling and Evaluating Sparse Autoencoders / Cunningham et al. 2023 (sparse autoencoders) Train an overcomplete autoencoder with a sparsity penalty on model activations, so that individual learned features stand for interpretable directions.
Findings IC-003 HyperDAS dynamically selects intervention tokens and learns linear subspaces in Llama3-8b that mediate entity attributes. [compared-to] IC-012 Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. [primary] IC-014 Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses. [primary] IC-060 SAE features in Pythia-160m and Mamba-130m exhibit high cross-architecture similarity with a depth-scaled correspondence [primary] IC-062 Llama 3 70B implements temporal difference learning in-context for reward-based RL, with causally relevant SAE features in its residual stream, while Llama 3 8B performs at chance [primary] IC-063 Llama 3 70B learns global graph structure via TD learning, building successor-representation-like geometry in its residual stream that is causally supported by TD latents [primary] IC-064 The TD learning mechanism identified in Llama 3 70B generalizes to Gemma-2-27B and Qwen-2.5-72B across all three tasks [primary] IC-1370 MLP0 representations of ordinal-sequence tokens in Pythia-1.4b contain linearly decodable mod-10 features that are causally important for incrementation [builds-on] IC-1370 MLP0 representations of ordinal-sequence tokens in Pythia-1.4b contain linearly decodable mod-10 features that are causally important for incrementation [primary] IC-296 The degree to which SAE features are active at multiple residual-stream layers increases with model size in Pythia, Gemma 2, Llama 3.2, and GPT-2 [builds-on] IC-296 The degree to which SAE features are active at multiple residual-stream layers increases with model size in Pythia, Gemma 2, Llama 3.2, and GPT-2 [supporting] IC-358 GPT-2-small and Mistral 7B contain circular representations of days of the week and months of the year in their internal activations, discovered via SAE dictionary element clustering [primary] IC-359 Mistral 7B and Llama 3 8B causally use circular subspaces to compute modular arithmetic on days of the week and months of the year [supporting] IC-525 GPT-2 small's residual stream at layer 8 decomposes into two sub-spaces of approximately 25% and 75% of the dimensionality [supporting] IC-526 GPT-2 small's first token position has residual stream norms more than an order of magnitude larger than all other positions [supporting] IC-527 A 16 million latent sparse autoencoder substituted into GPT-4 yields a language modeling loss corresponding to 10% of GPT-4's pretraining compute [primary] IC-808 Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layers [primary] IC-809 Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergence [primary] IC-810 Individual sparse autoencoder features in Pythia-70m-deduped are monosemantic and have predictable causal effects on output logits, as demonstrated by an apostrophe feature whose ablation primarily suppresses the 's' token [primary]