anchor
Findings
- IC-012Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. [source]
- IC-043Five ~7B decoder-only LLMs develop a high-intrinsic-dimensionality phase in intermediate layers that marks the transition from surface-form to abstract linguistic processing, with earlier onset predicting better next-token prediction [eval]
- IC-060SAE features in Pythia-160m and Mamba-130m exhibit high cross-architecture similarity with a depth-scaled correspondence [source]
- IC-135Linear relational embeddings for factual relations form in OLMo-7B, OLMo-1B, and GPT-J when subject-object co-occurrence frequency exceeds model-specific thresholds, with r=0.82 correlation between log co-occurrence and causality across all pretraining stages [source]
- IC-136LRE quality metrics from OLMo-7B predict pretraining term frequencies in GPT-J (trained on different data) with approximately 70% within-magnitude accuracy for object frequencies, outperforming log-probability-only features by about 30% [source]
- IC-185In Pythia-1B and Amber-7B, the probability of memorizing a training sequence scales log-linearly with both the number of repetitions in the corpus and the z-complexity of the sequence [source]
- IC-186The memorization status of sequences in Pythia-1B and Amber-7B is stationary throughout training: KL-LD fluctuations are mean-reverting with fixed variance, rejecting a random-walk model with p < 10⁻⁸ [source]
- IC-187Latent memorized sequences in Pythia-1B and Amber-7B can be recovered by adding random Gaussian noise of magnitude 2×10⁻³ to model weights, while un-memorized and unseen sequences cannot [source]
- IC-249The GitHub data-refined LLC identifies the induction circuit heads in Pythia-70m by distinguishing previous-token and induction heads from other head types across layers 2 and 3 [eval]
- IC-296The degree to which SAE features are active at multiple residual-stream layers increases with model size in Pythia, Gemma 2, Llama 3.2, and GPT-2 [source]
- IC-297Applying tuned-lens transformations to the residual stream decreases the apparent multi-layer SAE feature activity from 54–88% to 37–41% of total variance [source]
- IC-358GPT-2-small and Mistral 7B contain circular representations of days of the week and months of the year in their internal activations, discovered via SAE dictionary element clustering [source]
- IC-808Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layers [source]
- IC-809Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergence [source]
- IC-810Individual sparse autoencoder features in Pythia-70m-deduped are monosemantic and have predictable causal effects on output logits, as demonstrated by an apostrophe feature whose ablation primarily suppresses the 's' token [source]