Light Dark Distances between points in the representation track distances in the world the inputs came from, so that traversing the representation costs what traversing reality costs. Stronger than recovering a coordinate, because it is a claim about the whole geometry rather than about one readable direction.
Findings IC-063 Llama 3 70B learns global graph structure via TD learning, building successor-representation-like geometry in its residual stream that is causally supported by TD latents IC-108 Per-patch logit lens confidence in LLaVA localizes objects spatially, achieving mAP 79.90 on ImageNet segmentation, 8.03% above raw VLM attention IC-123 Llama-2-7B, Gemma-7B, and Llama-2-13B organize 16 concepts into hierarchical clusters in their representation space that reflect real-world category structure IC-125 LanguageBind's direct evaluation achieves 70% recall@10 on AudioSet, empirically validating that the inner product between unpaired modality representations recovers the correct probability ratio IC-127 ImageBind's direct evaluation closely matches logsumexp for both vision-language and audio-language alignment, validating the law for ImageBind IC-1283 CLIP's InfoNCE training objective is mathematically equivalent to performing generalized spectral clustering on the bipartite image-text pair graph IC-137 Pre-trained ResNet34 and ViT-B features on CIFAR-100 exhibit a block-diagonal class-correlation structure, with ViT-B showing higher intra-class correlation (0.35) than ResNet34 (0.25) IC-154 CLIP ViT-L/14 text embeddings fail to capture fine-grained visual class similarities, ranking rottweiler and doberman at position 828 behind unrelated pairs IC-158 Fine-tuning LLaVA-Med on MediConfusion training pairs cannot achieve 100% training accuracy, indicating the vision encoder's embeddings are fundamentally ambiguous for the confusing pairs IC-1619 DINOv2's patch-level features outperform CLIP and MAE for cross-image semantic feature matching IC-1630 Llama and Pythia models represent entity-attribute bindings via additive binding id vectors that form a continuous subspace with metric structure IC-280 CLIP, OpenCLIP, and SigLIP exhibit intra-modal misalignment: intra-modal similarity comparisons are suboptimal for image-to-image and text-to-text retrieval IC-281 SLIP's intra-modal self-supervised loss reduces intra-modal misalignment, making inter-modal inversion unnecessary for image retrieval IC-292 Neuron paths in ViT-B/16 show class-specific neuron clustering and semantic similarity between image categories IC-357 Off-the-shelf foundation models (DINO, CLIP, DINOv2, ViT) exhibit higher variance in their cosine similarity distributions than dataset-specific models, reducing the discriminative power of cosine similarity retrieval IC-481 Llama-3.1-8B and four other released LLMs reorganize their internal representations to reflect in-context graph structure in a sudden two-phase transition as context length increases IC-984 CLIP ViT-B/16's representation space does not reliably preserve semantic similarity as measured by shared image tags TM-014 Per-edge latent distance along shortest paths spikes at physical barriers