Light Dark A property of the input is encoded so that it can be recovered by a linear operation on the model's internal activations: a direction, a plane, or a linear read-out. Recoverability alone does not show the model uses the encoding for anything, which is why findings of this kind split into observational ones and interventional ones.
Findings IC-003 HyperDAS dynamically selects intervention tokens and learns linear subspaces in Llama3-8b that mediate entity attributes. IC-012 Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. IC-013 Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models. IC-014 Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses. IC-019 Visual token representations in LLaVA-1.5 evolve to align with interpretable text tokens IC-038 Few embedding dimensions drive the modality gap in CLIP and SigLIP IC-044 Tulu-2-13B's internal activations contain a linearly decodable, faithful representation of input-context propositions that persists under prompt injection and backdoor attacks where outputs become unfaithful IC-045 A 50-dimensional Hessian-identified subspace in Tulu-2-13B causally mediates entity-attribute binding, generalizing to three-entity contexts IC-052 GPT-2 small's IOI circuit activations are linearly decomposable into features for the io, s, and pos attributes, with the l10h0 name mover's attention decomposing into sparse pairwise feature interactions IC-053 In GPT-2 small's l10h0 name mover queries, the io attribute is encoded with higher-magnitude features than the s attribute, and both are causally relevant, but SAEs preferentially learn io features due to the magnitude asymmetry IC-060 SAE features in Pythia-160m and Mamba-130m exhibit high cross-architecture similarity with a depth-scaled correspondence IC-062 Llama 3 70B implements temporal difference learning in-context for reward-based RL, with causally relevant SAE features in its residual stream, while Llama 3 8B performs at chance IC-063 Llama 3 70B learns global graph structure via TD learning, building successor-representation-like geometry in its residual stream that is causally supported by TD latents IC-064 The TD learning mechanism identified in Llama 3 70B generalizes to Gemma-2-27B and Qwen-2.5-72B across all three tasks IC-106 Logit lens on LLaVA and InstructBLIP image representations shows higher internal confidence for objects present in the image than for hallucinated objects IC-107 Linear orthogonalization of LLaVA and InstructBLIP image features against text embeddings removes hallucinated objects at 83-86% individual rate versus 7-16% for correctly detected objects IC-1127 The LM head in GPT-2, GPT-J, BLOOM, Pythia, and LLaMA-2 projects all input token hidden states into interpretable token distributions over the vocabulary, and these distributions converge approximately monotonically toward the final layer's distribution IC-1201 CLIP-ViT-L/14 image features support 200-way zero-shot EEG-based object recognition better than ViT-B/16 or ResNet-50 features when used as a frozen encoder in a contrastive learning framework IC-1226 SAM's ViT-B encoder achieves 54.2% ImageNet-1k linear probing accuracy versus 67.7% for MAE's ViT-B, indicating its segmentation pretraining impairs high-level semantic representation IC-123 Llama-2-7B, Gemma-7B, and Llama-2-13B organize 16 concepts into hierarchical clusters in their representation space that reflect real-world category structure IC-125 LanguageBind's direct evaluation achieves 70% recall@10 on AudioSet, empirically validating that the inner product between unpaired modality representations recovers the correct probability ratio IC-1261 LLaMA-2 attention to constraint tokens correlates with factual correctness, and a linear probe on these attention weights predicts factual errors comparably to model confidence IC-127 ImageBind's direct evaluation closely matches logsumexp for both vision-language and audio-language alignment, validating the law for ImageBind IC-1317 Llama-2 and Pythia models contain linear representations of space and time that improve with depth and model scale IC-1318 Individual space and time neurons in Llama-2-7B causally contribute to spatial and temporal predictions IC-135 Linear relational embeddings for factual relations form in OLMo-7B, OLMo-1B, and GPT-J when subject-object co-occurrence frequency exceeds model-specific thresholds, with r=0.82 correlation between log co-occurrence and causality across all pretraining stages IC-136 LRE quality metrics from OLMo-7B predict pretraining term frequencies in GPT-J (trained on different data) with approximately 70% within-magnitude accuracy for object frequencies, outperforming log-probability-only features by about 30% IC-149 Intervening in the shared representation space using the dominant language (English) predictably changes model outputs for other data types, demonstrating the space is causally used rather than a vestigial byproduct IC-151 Each CLIP neuron's second-order effect is approximately a single linear direction in the joint text-image space, significant for fewer than 2% of images IC-1553 GPT-J, GPT-2-XL, and Llama-13B decode approximately 48% of tested relations via a linear transformation on the subject representation, and this structure causally influences predictions IC-160 Pythia-70m and Gemma-2-2b implement subject-verb agreement across a relative clause via a circuit of number detectors, PP/RC boundary detectors, and verb form promoters, with Gemma-2-2b additionally using NP number trackers IC-1630 Llama and Pythia models represent entity-attribute bindings via additive binding id vectors that form a continuous subspace with metric structure IC-166 A 1-dimensional subspace in a single layer encodes the context-versus-prior decision in Llama-3.1-8B, Gemma-2 9B, and Mistral-v0.3 7B, and setting this subspace steers the released (non-fine-tuned) models' behavior IC-167 Adding a PCA-derived control vector to the middle-layer residual stream improves logit-based reasoning accuracy on Pythia-1.4b, Pythia-2.8b, and Mistral-7B-Instruct IC-168 Control vectors derived from BABI improve GSM8K accuracy and vice versa on Mistral-7B-Instruct, indicating a task-general reasoning direction in the residual stream IC-181 Truthfulness in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct is linearly decodable from internal representations at exact answer tokens, with middle-to-late layers being most informative IC-182 Truthfulness encoding in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct is skill-specific rather than universal; probing classifiers do not meaningfully generalize across different task types beyond logit-based baselines IC-183 Error types in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct are linearly predictable from internal representations, encoding fine-grained information beyond binary correctness IC-184 Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct can internally encode the correct answer while externally generating an incorrect one, with the discrepancy most pronounced for error types where the model shows no external preference for the correct answer IC-246 Pretrained LLMs assign coherent semantic meaning to linear interpolations between token embeddings, extending the linear embedding hypothesis to the output space IC-346 Hierarchical and categorical concepts from WordNet are linearly represented in the final-layer space of Gemma-2b and Llama-3-8B, with semantic hierarchy encoded as orthogonality and categorical concepts as polytopes IC-413 Factors of variation in ImageNet-X are linearly decodable from the second-to-last-layer representations of ImageNet-pretrained ResNet50 and ViT-B/16 IC-419 CLIP ViT-B/16's layer-11 residual stream contains class-discriminative information in sparse SAE latent directions, and ablating class-specific top-k latents significantly degrades zero-shot classification accuracy IC-430 Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal IC-462 GPT-2 encodes toxicity in a low-dimensional linear subspace of its MLP layers, concentrated in higher layers IC-463 DPO's first-step gradients in GPT-2 are correlated with the toxic subspace, with stronger alignment in later layers and with more samples IC-468 Llama-3.1-405B's last hidden layer embeddings of erroneous tokens contain a linearly detectable error signal that a simple logistic regression head can exploit to flag incorrect continuations IC-481 Llama-3.1-8B and four other released LLMs reorganize their internal representations to reflect in-context graph structure in a sudden two-phase transition as context length increases IC-483 When in-context graph structure conflicts with pretrained semantic priors, Llama-3.1-8B encodes the in-context structure in higher principal components while the semantic prior dominates the first two IC-484 Re-scaling the first 2-3 principal components of Llama-3.1-8B token representations causally shifts next-token predictions toward the target graph position IC-493 A linear direction in the input embedding space of Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-Instruct-v0.3, and Phi-3-mini-128k predicts instruction-following success, generalizes across tasks but not instruction types, and can be used to improve adherence via representation engineering IC-494 The instruction-following dimension in Llama-2-7B-Chat and Llama-2-13B-Chat is more closely aligned with prompt phrasing than with task familiarity or instruction difficulty IC-499 ViT patch embeddings contain local semantic information beyond the [cls] token, as shown by performance degradation when restricting the output head to [cls] only or removing positional embeddings IC-501 Linear probes on middle-layer attention heads of Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 predict US lawmakers' DW-Nominate ideology scores with Spearman correlations around 0.85 IC-502 Linear probes trained on US lawmaker ideology generalize to predict Ad Fontes media slant scores when the same models simulate news outlets IC-503 Adding probe regression coefficients to attention head activations steers Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 toward more liberal or conservative generated text IC-505 Adversarial attacks on Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT shift hidden representations along the negative refusal feature direction IC-506 Restoring the refusal feature in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT causally disables all four tested adversarial attacks IC-507 The refusal feature direction in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT ranks near the top among 100 perturbations for compromising model safety IC-553 OpenCLIP ViT-B/16's image-text alignment score is a strong predictor of domain generalization accuracy, while perceptual similarity to LAION-400M pre-training data is a weaker predictor IC-561 A linear direction in the residual stream at layer 16 of Llama3-8B-Instruct is causally necessary and sufficient for self-authorship claims: steering with it achieves 100% control over authorship assertions, and projecting it out reduces claims by 50-60%. IC-562 Applying the layer-16 self-recognition vector to input tokens (not output) of Llama3-8B-Instruct alters the model's perception of authorship, making it believe or disbelieve it wrote arbitrary texts in both individual and paired paradigms. IC-565 Phi-3's residual stream encodes format instructions as linear directions, evidenced by cosine similarity and vocabulary-space projections IC-566 Adding instruction-specific steering vectors to the residual stream improves instruction-following accuracy for Phi-3, Gemma 2 2B IT, Mistral 7B IT, and Gemma 2 9B IT across format, length, and word-specific constraints IC-567 Steering vectors computed on instruction-tuned Gemma 2 models transfer to base Gemma 2 models, with cross-model steering outperforming same-model steering for Gemma 2 2B IC-568 In Phi-3, word-exclusion steering vectors computed via difference-in-means project onto the vocabulary space with high logits for the excluded word, making them counterproductive IC-578 LLMs encode input text as linearly separable representations in forerunner token hidden states, emerging in early layers and enhanced by in-context demonstrations IC-581 Induction heads for ICL operate on task-specific attention subspaces, with partial overlap across tasks, and the geometry of these subspaces explains demonstration saturation IC-593 The L2 norms of attention heads in Mistral-7B-Instruct and LLaMA-2-7B correlate with truthfulness, spiking by up to 83% at token positions of factual proposition completions and pertinent factual associations, and this correlation is specific to multi-headed attention representations rather than query, key, value, output, or FFN norms. IC-604 CLIP ViT-B/32's CIFAR-10 image embeddings approximately satisfy a multi-cluster structure with near-orthogonal class-mean features IC-645 CLIP, PickScore, and HPSv2 text embeddings share a common direction (cone effect) that captures text-irrelevant preferences, and the orthogonal component c⊥p better measures T2I alignment; CLIP's untrained common direction makes it ineffective for reward fine-tuning IC-678 Specific attention heads in CLIP ViT-L's last 4 layers encode specific image properties (color, shape, location, counting, texture) that are linearly recoverable via text directions IC-715 Factual information deleted from GPT-J, LLaMA-2, and GPT-2-XL via ROME or MEMIT remains linearly recoverable from intermediate hidden states, with up to 89% extraction success at budget b=20 IC-745 Function vectors for simple list-oriented tasks can be algebraically combined via addition and subtraction to produce new vectors that trigger composed tasks, sometimes outperforming 10-shot ICL. IC-796 GPT-2 Small MLP weight matrices are full-rank across all 12 layers and residual stream features are linearly recoverable from post-GELU MLP hidden activations, providing the structural conditions for the subspace patching illusion IC-808 Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layers IC-810 Individual sparse autoencoder features in Pythia-70m-deduped are monosemantic and have predictable causal effects on output logits, as demonstrated by an apostrophe feature whose ablation primarily suppresses the 's' token TM-001 A mean-difference vector between before and after image pairs acts as a transferable concept vector TM-002 The mean-difference concept vector scores higher than trained classifiers under cosine similarity TM-006 Coordinates are recoverable from TerraMind's frozen features, latitude more accurately than longitude TM-007 Coordinates are linearly decodable only in the larger TerraMind variants TM-013 Intervening on the altitude plane raises generated terrain by almost 500 metres