IC-358GPT-2-small and Mistral 7B contain circular representations of days of the week and months of the year in their internal activations, discovered via SAE dictionary element clustering

Joshua Engels, Eric J Michaud, Isaac Liao, Wes Gurnee, Max Tegmark

SourceNot All Language Model Features Are One-Dimensionally Linear

The authors train or obtain sparse autoencoders on GPT-2-small (all layers, from Bloom 2024) and Mistral 7B (layers 8, 16, 24, trained by the authors on 1B+ tokens). By clustering SAE dictionary elements by cosine similarity and reconstructing activations within each cluster, they find that days of the week, months of the year, and (in GPT-2) years of the 20th century are arranged in circular order in the PCA projections of the reconstructed activations. The days-of-week cluster in GPT-2-small scores m_ε(f)=0.4750 and s(f)=0.9506 on the irreducibility tests, ranking 9th out of 1000 clusters by the combined metric.

Evidence
observational
Key metric
m_ε(f) = 0.4750, s(f) = 0.9506 for days-of-week cluster (GPT-2-small, layer 7); clusters rank 9, 28, 15 out of 1000 by product of (1 − ε-mixture index) and separability index
Caveat
The authors note it is unclear why they did not find more interpretable multi-dimensional features, and their irreducibility definitions are purely statistical and had to be relaxed to hold in practice.
Model
GPT-2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
Concepts
Circular representation
Datasets
The Pile [source]
Methods
Sparse autoencoder / Sparse autoencoders / K-sparse autoencoder / Topk sparse autoencoder / Scaling and Evaluating Sparse Autoencoders / Cunningham et al. 2023 (sparse autoencoders) [primary], Spectral Clustering [supporting], Principal component analysis [supporting]
Related findings
IC-359, IC-360, IC-361
Extraction
automatic-extraction