Light Dark Principal component analysis anchor
Rotates the data onto orthogonal directions of decreasing variance. Used both to summarise how concentrated a representation is and, on the first two components, to draw it. The share of variance the top components carry is a claim about the representation; the picture is not.
Findings IC-123 Llama-2-7B, Gemma-7B, and Llama-2-13B organize 16 concepts into hierarchical clusters in their representation space that reflect real-world category structure [primary] IC-1610 Llama-2-7b-chat underperforms on small molecule editing tasks due to limited domain-specific pretraining [compared-to] IC-167 Adding a PCA-derived control vector to the middle-layer residual stream improves logit-based reasoning accuracy on Pythia-1.4b, Pythia-2.8b, and Mistral-7B-Instruct [primary] IC-168 Control vectors derived from BABI improve GSM8K accuracy and vice versa on Mistral-7B-Instruct, indicating a task-general reasoning direction in the residual stream [primary] IC-357 Off-the-shelf foundation models (DINO, CLIP, DINOv2, ViT) exhibit higher variance in their cosine similarity distributions than dataset-specific models, reducing the discriminative power of cosine similarity retrieval [compared-to] IC-358 GPT-2-small and Mistral 7B contain circular representations of days of the week and months of the year in their internal activations, discovered via SAE dictionary element clustering [supporting] IC-361 Mistral 7B's circular representation of days of the week is continuous, mapping intermediate time-of-day values to positions between adjacent weekdays [primary] IC-481 Llama-3.1-8B and four other released LLMs reorganize their internal representations to reflect in-context graph structure in a sudden two-phase transition as context length increases [primary] IC-482 In Llama-3.1-8B, in-context graph structure is absent in early layers dominated by semantic priors and emerges clearly in deeper layers [primary] IC-483 When in-context graph structure conflicts with pretrained semantic priors, Llama-3.1-8B encodes the in-context structure in higher principal components while the semantic prior dominates the first two [primary] IC-484 Re-scaling the first 2-3 principal components of Llama-3.1-8B token representations causally shifts next-token predictions toward the target graph position [primary] IC-493 A linear direction in the input embedding space of Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-Instruct-v0.3, and Phi-3-mini-128k predicts instruction-following success, generalizes across tasks but not instruction types, and can be used to improve adherence via representation engineering [supporting] IC-581 Induction heads for ICL operate on task-specific attention subspaces, with partial overlap across tasks, and the geometry of these subspaces explains demonstration saturation [supporting] IC-582 Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIs [supporting] IC-673 Trained depthwise convolutional kernels in DS-CNN architectures converge to identifiable DoG-like patterns, with over 95% of ConvNeXtV2 and over 90% of ConvNeXt filters classifiable into a small set of clusters [supporting] IC-808 Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layers [compared-to] IC-809 Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergence [compared-to] IC-992 LLaMA-2-7B's last hidden layer encodes the prompt format with high identifiability, and the separability of format embeddings in the top two principal components correlates with performance spread [supporting] TM-008 TerraMind embeddings separate hemispheres rather than climate zones [supporting] TM-009 A weak seasonal shift in the embeddings matches climatic intuition but is attributed to geography [supporting] TM-014 Per-edge latent distance along shortest paths spikes at physical barriers [supporting]