- FX-001Pairwise Banzhaf interactions explain CLIP similarity more faithfully than single-score methodsCLIP / CLIP-ViT (LC)
- FX-002FIXLIP gives SigLIP-2 higher pointing-game recognition than CLIP at ViT-B/32 and ViT-B/16CLIP / CLIP-ViT (LC), SigLIP, SigLIP-2
- FX-003FIXLIP's strongest interaction in one CLIP example links doll to an image patch reading dollarCLIP / CLIP-ViT (LC)
- IC-001Vision-language models perform near chance on the NL-Eye visual abductive reasoning benchmarkGemini 1.5 / Gemini Pro 1.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Claude 3.5, Claude 3, LLaVA-NeXT / LLaVA 1.6, Fuyu, MiniCPM-V, LLaVA-OneVision
- IC-002Even when VLMs select the correct hypothesis, their explanations are often invalid or unhelpfulGemini 1.5 / Gemini Pro 1.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Claude 3.5, Claude 3, LLaVA-NeXT / LLaVA 1.6, Fuyu
- IC-003HyperDAS dynamically selects intervention tokens and learns linear subspaces in Llama3-8b that mediate entity attributes.Llama 3
- IC-004Retrieval heads are sparse, universal, and causally responsible for long-context retrieval in LLMsLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Yi, Qwen1.5, Jamba
- IC-005GPT-4o underperforms AHA and other VLMs in detecting and reasoning about robotic manipulation failures across multiple datasets.GPT-4o
- IC-006The Retriever-Dictionary module improves object detection accuracy of YOLOv7, YOLOv9, Faster R-CNN, and Deformable DETR on COCO 2017YOLOv7, YOLOv9, Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101, Deformable DETR
- IC-007Most LLMs do not align closely with human moral preferences on multilingual trolley problemsGPT-3 / GPT base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Llama 2 / Llama 2 base, Llama 3, Llama 3.1, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Phi-3, Qwen 2
- IC-008Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuancesGPT-4o, Llama 3.1
- IC-009LLM moral preferences show significant language sensitivity but not inequality toward low-resource languagesGPT-3 / GPT base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Llama 2 / Llama 2 base, Llama 3, Llama 3.1, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Phi-3, Qwen 2
- IC-010LLM responses to trolley problems are moderately robust across prompt paraphrasesLlama 3
- IC-011Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferencesLlama 3.1, Gemma 2, Qwen 2, Llama 2 / Llama 2 base
- IC-012Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal.Gemma 2, Llama 3.1
- IC-013Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models.Gemma 2, Llama 3.1
- IC-014Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses.Gemma 2
- IC-015Truncating MLP weights in Pythia-1b increases the probability of the correct answer in an Indirect Object Identification taskPythia
- IC-016Truncating MLP weights in Pythia-1b increases the probability of the correct answer in a factual recall taskPythia
- IC-017Truncating MLP weights improves few-shot Chain-of-Thought reasoning accuracy on GSM8K for Phi-3 and Llama-3.1-8BPhi-3, Llama 3.1
- IC-018Object information is localized to specific visual tokens in LLaVA-1.5LLaVA-1.5 / LLaVA-v1.5, LLaVA-Phi
- IC-019Visual token representations in LLaVA-1.5 evolve to align with interpretable text tokensLLaVA-1.5 / LLaVA-v1.5, Qwen2-VL
- IC-020LLaVA-1.5 extracts object information directly from visual tokens to the last token in mid-late layersLLaVA-1.5 / LLaVA-v1.5, LLaVA-Phi
- IC-021Vision-language adaptation degrades safety in Llama-2-chat-7b even when training data is filtered for safetyLlama 2 / Llama 2 base
- IC-022Safety layers in Llama-2-chat-7b show substantial divergence during VL adaptation, correlating with safety degradationLlama 2 / Llama 2 base
- IC-023Local scaling, rank, and complexity of Stable Diffusion correlate with generation aesthetics, diversity, and memorizationStable Diffusion
- IC-024Reward model trained on local scaling of Stable Diffusion can guide generation to increase diversity and aesthetic scoresStable Diffusion
- IC-025LMMs exhibit poor fine-grained perception in locating individual characters on original oracle bonesGemini 1.5 / Gemini Pro 1.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Qwen-VL, XGen-MM, mPLUG-Owl3, MiniCPM-V, Moondream2, InternVL2, GLM-4V, CogVLM2, LLaVA-NeXT / LLaVA 1.6, Idefics2, DeepSeek-VL, InternLM-XComposer2-VL, LLaVA-1.5 / LLaVA-v1.5
- IC-026LLMs can assist in OB rejoining by identifying rejoinable fragments with moderate accuracy, but are not yet truly usableGemini 1.5 / Gemini Pro 1.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Qwen-VL, GLM-4V, XGen-MM, mPLUG-Owl3, MiniCPM-V, InternVL2, LLaVA-NeXT / LLaVA 1.6, Idefics2, DeepSeek-VL
- IC-027LMM performance in deciphering oracle bone inscriptions is comparable to untrained humans for common characters but declines for rarer and structurally complex charactersGemini 1.5 / Gemini Pro 1.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Qwen-VL, XGen-MM, mPLUG-Owl3, MiniCPM-V, Moondream2, InternVL2, GLM-4V, CogVLM2, LLaVA-NeXT / LLaVA 1.6, Idefics2, DeepSeek-VL, InternLM-XComposer2-VL, LLaVA-1.5 / LLaVA-v1.5
- IC-028SPADE, an abstaining classifier built on top of ResNet, ViT, and VGG models, detects out-of-distribution and adversarial samples with provable guarantees.ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, VGG / VGG13, ViT
- IC-029Large language models show conformity to group answers in multi-agent interactionsGPT-3.5 / ChatGPT-3.5, GPT-4o, Llama 3, Llama 3.1, Gemma 2, Qwen 2, GLM-4
- IC-030Larger language models exhibit higher independence rates and lower conformity under some protocolsQwen 2, Llama 3, Llama 3.1, Gemma 2, GPT-4o, GPT-3.5 / ChatGPT-3.5
- IC-031Empowered persona prompts and reflection mechanisms reduce conformity in large language modelsLlama 3, Qwen 2
- IC-032Off-policy DPO causes a squeezing effect in LLMs where probability mass shifts to the most confident token, explaining degenerate repetitionPythia, Qwen1.5
- IC-033Pre-training the SFT stage on both chosen and rejected responses mitigates the DPO squeezing effect and improves alignment win ratesQwen1.5
- IC-034Benefit and detriment in RAG can be traded off at token level for Llama-2, OPT and Mistral using representation similarityOPT, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-035Removing the inductive bias of locality from Vision Transformers improves or matches performance on classification and regression tasks.ViT
- IC-036Removing locality from Vision Transformers improves performance in self-supervised learning via Masked Autoencoding.ViT
- IC-037Removing locality from Diffusion Transformers improves image generation quality.DiT
- IC-038Few embedding dimensions drive the modality gap in CLIP and SigLIPCLIP / CLIP-ViT (LC), SigLIP
- IC-039Object bias in CLIP and SigLIP is not correlated with performance on attribute tasksCLIP / CLIP-ViT (LC), SigLIP
- IC-040Information imbalance triggers both the modality gap and object bias in contrastive VLMsCLIP / CLIP-ViT (LC)
- IC-041CLIP and SigLIP use the modality gap to control logit entropyCLIP / CLIP-ViT (LC)
- IC-042No single knowledge editing method excels across all criteria when editing visual and user-specific knowledge in LMMs.BLIP-2, MiniGPT-4, LLaVA-1.5 / LLaVA-v1.5
- IC-043Five ~7B decoder-only LLMs develop a high-intrinsic-dimensionality phase in intermediate layers that marks the transition from surface-form to abstract linguistic processing, with earlier onset predicting better next-token predictionOPT, Llama 3, Pythia, OLMo / OLMo base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-044Tulu-2-13B's internal activations contain a linearly decodable, faithful representation of input-context propositions that persists under prompt injection and backdoor attacks where outputs become unfaithfulTulu 2, Llama 2 / Llama 2 base
- IC-045A 50-dimensional Hessian-identified subspace in Tulu-2-13B causally mediates entity-attribute binding, generalizing to three-entity contextsTulu 2, Llama 2 / Llama 2 base
- IC-046Tulu-2-13B exhibits gender bias in both its internal binding representation and its outputs, with the output-level bias being stronger than the representation-level biasTulu 2, Llama 2 / Llama 2 base
- IC-047Tulu-2-13B's entity-attribute binding partially relies on token order as a shortcut, degrading in nested orderings where order and semantic binding conflictTulu 2, Llama 2 / Llama 2 base
- IC-048Existing video LLMs (TimeChat, VTG-LLM, Momentor, Hawkeye) show limited zero-shot video temporal grounding capability and struggle to improve with fine-tuningTimeChat, VTG-LLM, Momentor, Hawkeye
- IC-049GPT-4o shows strong performance on some E.T.Bench event-level tasks (RVQ: 57.7, VHD: 56.9) but very weak performance on others (EPM: 4.5, TAL: 20.0)GPT-4o
- IC-050GPT-4o achieves 53.33 overall on Event-Bench with strong event description (57.50) and counter reasoning (63.44) but weaker episodic reasoning (37.33)GPT-4o
- IC-051Qwen2-VL (7B) achieves 0.0 on DVC, DVC SLC, and TEM tasks on E.T.BenchQwen2-VL
- IC-052GPT-2 small's IOI circuit activations are linearly decomposable into features for the io, s, and pos attributes, with the l10h0 name mover's attention decomposing into sparse pairwise feature interactionsGPT-2
- IC-053In GPT-2 small's l10h0 name mover queries, the io attribute is encoded with higher-magnitude features than the s attribute, and both are causally relevant, but SAEs preferentially learn io features due to the magnitude asymmetryGPT-2
- IC-054Gemma 2 2B performance degrades substantially when routed through Gemma Scope SAEs, and SAE-based feature suppression causes broad cross-domain degradation rather than targeted knowledge removalGemma 2
- IC-055OLMoE 6.9B experts show no domain specialization, with routing scores evenly distributed across MMLU domains, preventing targeted knowledge unlearningOLMoE 6.9B
- IC-056Influence scores of pretraining documents correlate across reasoning queries of the same type, indicating Command R 7B and 35B rely on shared procedural knowledge rather than retrieving specific answersCohere Command R
- IC-057Command R 7B and 35B rely on each individual pretraining document less per nat of generated information for reasoning than for factual questions, with less volatile influence magnitudesCohere Command R
- IC-058The answer to factual questions appears in the top 0.01% most influential pretraining documents for 55% of 7B queries and 30% of 35B queries, but almost never for reasoning questionsCohere Command R
- IC-059Code data is strongly overrepresented in the most influential pretraining documents for reasoning queries in Command R 7B and 35BCohere Command R
- IC-060SAE features in Pythia-160m and Mamba-130m exhibit high cross-architecture similarity with a depth-scaled correspondencePythia, Mamba
- IC-061The induction circuit in Mamba-130m is structurally analogous to the transformer induction circuit, with an off-by-one motif in SSM state writingMamba
- IC-062Llama 3 70B implements temporal difference learning in-context for reward-based RL, with causally relevant SAE features in its residual stream, while Llama 3 8B performs at chanceLlama 3
- IC-063Llama 3 70B learns global graph structure via TD learning, building successor-representation-like geometry in its residual stream that is causally supported by TD latentsLlama 3
- IC-064The TD learning mechanism identified in Llama 3 70B generalizes to Gemma-2-27B and Qwen-2.5-72B across all three tasksGemma 2, Qwen2.5
- IC-069EVA-CLIP's dense patch features are semantically contaminated by surrounding context, degrading their spatial qualityEVA-CLIP
- IC-070Region-language alignment fine-tuning degrades EVA-CLIP's spatial awareness as measured by unsupervised segmentationEVA-CLIP
- IC-071DINOv2's dense features are dominated by global context, impairing fine-grained spatial detailDINOv2
- IC-072LLM agents of varying scales exhibit a failure mode on web automation tasks when processing raw, complex web page observations, with the penalty being more severe for smaller modelsGPT-4o, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5, Llama 3.1
- IC-073Released LLMs (GPT-4o, Llama-3.1-70B, Qwen2-7B, etc.) show limited workflow orchestration capability that degrades as workflow complexity increasesGPT-4o, Llama 3.1, Qwen 2
- IC-074Released LLMs achieve F1 plan scores between 42.7 and 86.7 on the T-Eval plan taskGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Qwen, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3.1, Llama 2 / Llama 2 base, Vicuna, Baichuan2-13B, Wizardlm, Qwen1.5
- IC-075GPT-4o-mini and Qwen2.5-72B achieve low precision and recall when used as API retrievers for workflow orchestrationGPT-4o, Qwen2.5
- IC-076GoogLeNet and ViT exhibit input space mode connectivity: inputs with similar predictions are connected by low-loss paths, with real-real pairs showing approximately linear paths and real-adversarial pairs showing significantly higher barriersGoogLeNet, ViT
- IC-077VGG-16's loss landscape barrier height distinguishes adversarial from real inputs, enabling a detection method that outperforms baselines on DeepFool and C&W attacksVGG / VGG13
- IC-078GPT-4's self-verification loop causes performance collapse due to high false negative rates in binary verificationGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-079GPT-4's free-form critique generation is unreliable, containing hallucinated edges, vertex colors, and precondition statesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-080GPT-4's performance is largely insensitive to the content of feedback; simple re-prompting with a sound verifier (sampling) matches or exceeds detailed critiqueGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-081ViT/L-16 exhibits lower sensitivity to token-wise Gaussian perturbations than ConvNeXtV2-Tiny on ImageNet-1kViT, ConvNeXtV2-Tiny
- IC-082GPT-3.5, GPT-4o, Claude-3.5-Sonnet, and Llama-3.1-8B produce explanations on the BBQ social bias task that are systematically unfaithful for identity and behavior concepts while remaining faithful for context concepts, with specific patterns of hiding safety-measure influence and social biasGPT-3.5 / ChatGPT-3.5, GPT-4o, Claude 3.5, Llama 3.1
- IC-083GPT-3.5, GPT-4o, and Claude-3.5-Sonnet produce unfaithful explanations on MedQA medical questions, omitting high-effect clinical concepts such as the patient's mental status while over-referencing low-effect conceptsGPT-3.5 / ChatGPT-3.5, GPT-4o, Claude 3.5
- IC-084Safety alignment in Llama-2-7b-chat and Gemma-7b-1.1-it is shallow, with the KL divergence from the base model concentrated in the first few output tokens, making the models vulnerable to prefilling attacksLlama 2 / Llama 2 base, Gemma
- IC-085Unaligned base models Llama-2-7b and Gemma-7b produce predominantly safe continuations when prefilled with refusal prefixes, demonstrating a pre-existing safety shortcutLlama 2 / Llama 2 base, Gemma
- IC-086Fine-tuning Llama-2-7b-chat on 100 harmful examples for 6 gradient steps increases the attack success rate from 1.5% to 87.9%, with per-token dynamics showing the distributional change concentrated in the first few tokensLlama 2 / Llama 2 base
- IC-087Answer symbol production in OLMo 7B Instruct, Llama 3.1 8B Instruct, and Qwen 2.5 1.5B Instruct is causally attributed to a few middle layers and specifically their multi-head self-attention mechanisms, with a sparse set of 1-4 attention heads per layer responsibleOLMo / OLMo base, Llama 3.1, Qwen2.5
- IC-088OLMo 7B Instruct and Qwen 2.5 1.5B Instruct exhibit a two-stage process for unusual answer symbols, initially assigning non-negligible probability to expected symbols (a/b/c/d) before switching to the actual prompt symbols at a specific later layerOLMo / OLMo base, Qwen2.5
- IC-089OLMo 0724 7B base learns formatted multiple-choice question answering between 80k and 100k training steps, transitioning from near-random to near-perfect accuracy on the synthetic colors taskOLMo / OLMo base
- IC-092All evaluated LLMs show consistent F1 degradation to at most 0.60 when two or more events match a retrieval cueGPT-4o, Claude 3, Claude 3.5, Llama 3.1, O1 / OpenAI-o1-preview
- IC-093No evaluated LLM achieves perfect confabulation avoidance on questions about non-existent eventsGPT-4o, Claude 3, Claude 3.5, Llama 3.1, O1 / OpenAI-o1-preview
- IC-094Episodic recall accuracy degrades systematically from content cues to space cues to time cues across all evaluated LLMsGPT-4o, Claude 3, Claude 3.5, Llama 3.1, O1 / OpenAI-o1-preview
- IC-095Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological orderingGPT-4o, Claude 3, Claude 3.5, Llama 3.1, O1 / OpenAI-o1-preview
- IC-096Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the topLlama 3.1, Qwen 2.5 72B Instruct, O3, GPT-4o, Llama 3, GPT-4.5, DeepSeek R1, Claude 3.7 Sonnet
- IC-097Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3BLlama-3.2-3B, Llama 3.1
- IC-098LegalBERT performs below the constant classifier baseline on CASELAWQA due to its 512-token context windowLegalBERT, SAULLM 54B
- IC-099GPT-4 and Claude 3 Opus can be prompted to selectively underperform on WMDP while maintaining general performance on MMLU and CSQAGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3
- IC-100GPT-4, GPT-3.5, Claude 3, Llama 3 8B, and Llama 3 70B can be prompted to approximately calibrate their accuracy to specific target percentages on MMLUGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Claude 3, Llama 3
- IC-1003Pretrained ResNet-50 and ViT-B/16 exhibit neuron activation patterns that are separable between in-distribution and out-of-distribution inputs, enabling post-hoc OOD detection without model modificationResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, ViT
- IC-1005Social bias neurons in BERT-base-cased and RoBERTa-base are concentrated in the deepest transformer layersBERT, RoBERTa / RoBERTa-L
- IC-1007LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedbackChatGPT, GPT-3 / GPT base, Llama 2 / Llama 2 base
- IC-1008The magnitude of CRITIC's improvement on mathematical program synthesis scales with Llama-2 model sizeLlama 2 / Llama 2 base
- IC-101GPT-4 and Claude 3 struggle to emulate a lower capability profile (high school freshman level) via zero-shot prompting, with only moderate improvement from chain-of-thought promptingGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3
- IC-1011OpenAI CLIP loses approximately 8% zero-shot retrieval accuracy on 2021–2022 data compared to OpenCLIP models trained on data through 2022, while standard benchmarks show no such gapCLIP / CLIP-ViT (LC), OpenCLIP
- IC-1012Connected regions in the latent space of Stable Diffusion v2.1, v1.5, and GLIDE produce distorted images independent of the text promptStable Diffusion, GLIDE
- IC-1013Latent samples in Stable Diffusion v2.1, v1.5, and GLIDE can produce images of associated backgrounds rather than the key object, with failure rates of 2.7%, 9.2%, and 50.5% under random sampling respectivelyStable Diffusion, GLIDE
- IC-1014A single adversarial token embedding appended to any input prompt overwrites the prompt in Stable Diffusion v2.1 to generate a target object, with CLIP similarity to the original prompt (0.742) remaining higher than to the target (0.546)Stable Diffusion, GLIDE
- IC-1015GPT-J and 10 other LLMs exhibit overthinking: calibrated accuracy given incorrect few-shot demonstrations peaks at a critical layer then declines, and ablating 5 false induction heads in late layers reduces the accuracy gap by 38.9% on averageGPT-J, GPT-2, GPT-NeoX-20B, Pythia, Llama 2 / Llama 2 base
- IC-1016ImageNet-pretrained ResNet-50 backbone exhibits shortcut bias toward background features when adapted to bird classification via a new readout layerResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-1017Vision and language models pre-trained on noisy data exhibit degraded OOD transfer that is partially recoverable via SVD-based feature-space regularizationEfficientNet, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, Swin Transformer, ViT, ConvNeXt, BERT, RoBERTa / RoBERTa-L, GPT-2, text-ada-002
- IC-102GPT-4o's spatial understanding degrades when depth maps are provided as additional input on SpatialBenchGPT-4o
- IC-1023The binary activation pattern of standard CNNs (VGG, ResNet) carries most of the classification information, as shown by APoPVGG / VGG13, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-103LLMs' value rankings align with the universal human value hierarchy under most prompting conditionsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.0 Pro, Llama 3.1, Gemma 2
- IC-1035Reprover achieves 0% accuracy on sorry theorems in advanced mathematics repositories (PFR, Hairy Ball Theorem, Coxeter) while proving basic theorems in other repositoriesReprover
- IC-104Value anchor prompting produces LLM value correlation structures that closely match the human circular value structure, while standard prompting does notGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.0 Pro, Llama 3.1, Gemma 2
- IC-1044Counterfactual GNN explainers produce statistically infeasible recourses that violate topological constraints in molecular datasetsRCExplainer, CF2, CLEAR
- IC-1045Factual GNN explanations do not capture the full data signal: retraining on explanations fails to reproduce predictions while retraining on residuals preserves themPGExplainer, TAGExplainer, GEM, CF2, RCExplainer
- IC-105Value anchoring produces a sinusoidal scoring pattern around the circular value structure, with scores decreasing as circular distance from the anchor increasesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.0 Pro, Llama 3.1, Gemma 2
- IC-1050Released models generate patches that are less than half the length of gold solutions and rarely edit more than one fileGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1053Instruction tuning suppresses in-context learning in LLaMA, Vicuna, and OPT-IML, with the suppression being largest for English prompts and partially recoverable via translation to other languagesLLaMA, Alpaca, Vicuna, OPT
- IC-1054Code fine-tuning degrades English natural language reasoning in Code LLaMA relative to LLaMA-2, but the effect is negligible or slightly positive in French, Spanish, and GermanLlama 2 / Llama 2 base, Code Llama
- IC-1055Safety fine-tuning suppresses harmful content generation in ChatGPT relative to GPT-3.5, but the suppression is substantially weaker for non-English promptsGPT-3.5 / ChatGPT-3.5, ChatGPT
- IC-106Logit lens on LLaVA and InstructBLIP image representations shows higher internal confidence for objects present in the image than for hallucinated objectsLLaVA-1.5 / LLaVA-v1.5, InstructBLIP, LLaVA-NeXT / LLaVA 1.6, Cambrian-1
- IC-1069Stable Diffusion v2.1's compositional understanding on the ARO benchmark is significantly higher than previously reported by MMSE-based scoringStable Diffusion, OpenCLIP
- IC-107Linear orthogonalization of LLaVA and InstructBLIP image features against text embeddings removes hallucinated objects at 83-86% individual rate versus 7-16% for correctly detected objectsLLaVA-1.5 / LLaVA-v1.5, InstructBLIP, LLaVA-NeXT / LLaVA 1.6, Cambrian-1
- IC-1070In Stable Diffusion v2.1, attention maps do not reliably predict the effect of prompt interventions on generated images, while conditional mutual information doesStable Diffusion
- IC-1071Stable Diffusion v2.1's pixel-wise conditional mutual information localizes abstract words (adjectives, adverbs, verbs) more effectively than attention, but is less effective than attention for object segmentationStable Diffusion
- IC-1078LLaMA models (7B through 65B) exhibit gender bias in language generation, coreference resolution, and sentence likelihood, with stereotypical associations driving predictionsLLaMA
- IC-1079Causal tracing reveals that mid-upper MLP layers (layers 18–25 in 7B) are the primary mediators of stereotypical gender bias in LLaMA, while the last layers show negative coefficients that counter the biasLLaMA
- IC-108Per-patch logit lens confidence in LLaVA localizes objects spatially, achieving mAP 79.90 on ImageNet segmentation, 8.03% above raw VLM attentionLLaVA-1.5 / LLaVA-v1.5
- IC-1080OPUS-MT small, OPUS-MT large, and mBART-50 in their default (unfine-tuned) form achieve low context-sensitive disambiguation accuracy on discourse-level phenomenaOPUS-MT, mBART-50
- IC-1086RS LDS fails to increase its number of active states when the underlying dynamics change non-stationarilyRS-LDS
- IC-1087SLDS produces poor dynamical accuracy on the NASCAR task because it lacks recurrent switchingSLDS
- IC-1088ICL predictions in LLaMA, LLaMA-2, and Falcon models depend on in-context label information and can learn truly novel label relationshipsLlama 2 / Llama 2 base, LLaMA, Falcon
- IC-1089ICL in LLaMA, LLaMA-2, and Falcon models cannot fully overcome pre-training label preferences when in-context labels are flippedLlama 2 / Llama 2 base, LLaMA, Falcon
- IC-109GPT-3.5-turbo and GPT-4 produce cycles in inferred causal graphs when using pairwise prompts, with cycle counts growing sharply on larger graphsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1090ICL in LLaMA, LLaMA-2, and Falcon models preferentially uses in-context label information closer to the query rather than treating all examples equallyLlama 2 / Llama 2 base, LLaMA, Falcon
- IC-1099AdamW-pretrained vision models (ViTs, ConvNeXt) have disproportionately large embedding-layer gradients at initialization, causing SGD fine-tuning to degrade OOD accuracy by up to 15% relative to AdamWCLIP / CLIP-ViT (LC), ViT, DINO, ConvNeXt, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-110Phi-3 (3.8B) and Llama-3 (8B) with triplet prompting outperform GPT-4 with pairwise prompting on causal graph orientationPhi-3, Llama 3, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1105PAC-Bayes generalization bounds for discrete class prompts on CLIP are within a few percentage points of the actual test error across CIFAR-10, CIFAR-100, ImageNet, FMOW, and OfficeHomeCLIP / CLIP-ViT (LC)
- IC-1106CLIP prompts found by greedy search do not fit random labels: train and test error drop in tandem as the fraction of flipped labels increases, unlike a linear probe which achieves near-random accuracyCLIP / CLIP-ViT (LC)
- IC-111ViT-B/16 pretrained with MAE exhibits higher attention diversity than ViT-B/16 pretrained with MoCo v3, DINO, or DeiTViT
- IC-112Released LLMs show a reproducible failure mode where numerical task accuracy degrades sharply as input digit length increasesGPT-4o, Llama 3.1, Llama 2 / Llama 2 base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Qwen 2
- IC-1123All seven published concept erasure methods applied to Stable Diffusion 1.4 can be circumvented by learned word embeddings, demonstrating that targeted concepts are input-filtered rather than truly removed from the modelStable Diffusion
- IC-1127The LM head in GPT-2, GPT-J, BLOOM, Pythia, and LLaMA-2 projects all input token hidden states into interpretable token distributions over the vocabulary, and these distributions converge approximately monotonically toward the final layer's distributionGPT-2, GPT-J, BLOOM, Pythia, Llama 2 / Llama 2 base
- IC-113Released LLMs show a reproducible failure mode where accuracy on fraction and scientific notation tasks falls below 20% even for the shortest inputsGPT-4o, Llama 3.1, Llama 2 / Llama 2 base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Qwen 2
- IC-1133LLMs are highly receptive to coherent counter-memory as sole evidence, contradicting prior findings of stubbornness with entity-substitution counter-memoryChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, PaLM 2, Qwen, Llama 2 / Llama 2 base, Vicuna
- IC-1134LLMs show strong confirmation bias in multi-source settings, preferring evidence consistent with parametric memory, with stronger bias for popular entitiesChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, PaLM 2, Qwen, Llama 2 / Llama 2 base, Vicuna
- IC-1135LLMs show order sensitivity to evidence position in context, with PaLM2 and LLaMA2-7B showing memorization ratio variations exceeding 30%ChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, PaLM 2, Llama 2 / Llama 2 base
- IC-1136Larger LLMs (LLaMA2-70B, Vicuna-33B) are more stubborn than their smaller counterparts (LLaMA2-7B, Vicuna-7B) when encountering incoherent entity-substitution counter-memoryLlama 2 / Llama 2 base, Vicuna
- IC-1137Stable Diffusion represents certain concepts primarily through specific named exemplars rather than abstract category featuresStable Diffusion
- IC-1138Stable Diffusion simultaneously encodes multiple meanings of homograph concepts in a single representationStable Diffusion
- IC-1139Stable Diffusion encodes social biases in its internal concept representations that are not always visually apparentStable Diffusion
- IC-114Released LLMs cannot reliably identify a specific digit in a number as the number's length increases, with GPT-4o achieving only 20% on get-digit in the xl rangeGPT-4o, Llama 3.1, Llama 2 / Llama 2 base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Qwen 2
- IC-1140Stable Diffusion's internal concept representations encode visual and structural similarities (shape, texture, color) that transcend textual semanticsStable Diffusion
- IC-1144GPT-4 and GPT-3.5 produce high rates of irrelevant (fabricated) books when answering constraint queries from parametric knowledge, with a sharp phase transition at low author popularityGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1145Providing complete context eliminates irrelevance but does not fix constraint satisfaction for GPT-4 or GPT-3.5GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1146Self-context (self-retrieval) chain-of-thought increases the rate of fabricated books compared to no-context for both GPT-4 and GPT-3.5GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1147GPT-4 outperforms GPT-3.5 on all KITAB metrics but the gap is modest, with all-correctness below 35% for both, suggesting scale alone does not resolve constraint satisfactionGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1148ChatGPT and InstructGPT exhibit positional bias in their reliance on in-context examples, with ChatGPT showing decreasing attention by position and InstructGPT showing a U-shaped patternChatGPT, InstructGPT
- IC-1149ChatGPT and Llama-2-7b-chat fail to recognize unanswerable questions on SQuAD 2.0, with Llama-2-7b-chat scoring only 3.72% accuracy on no-answer questionsChatGPT, Llama 2 / Llama 2 base
- IC-115NUPA performance is largely independent of model size within a family: GPT-4o and GPT-4o-mini show nearly identical performance, as do Qwen2-72B and Qwen2-7BGPT-4o, Qwen 2
- IC-1150ChatGPT and Llama-2-7b-chat underperform humans by 20 and 31 points respectively on out-of-distribution NLU tasks in GLUE-XChatGPT, Llama 2 / Llama 2 base
- IC-1151LLaMA, OPT, LLaMA-2, Mistral, and GPT-J all exhibit token co-occurrence reinforcement, where the probability of generating a token increases monotonically with the number of its contextual co-occurrencesLLaMA, OPT, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, GPT-J
- IC-1152Token reinforcement in demonstrations constrains LLaMA-65B's output to valid label spaces on MMLU and enables chain-of-thought pattern following on GSM8K without requiring question contentLLaMA
- IC-1153Intentionally constructed spurious token connections in MMLU demonstrations misdirect LLaMA-65B's in-context learning toward specific answer choicesLLaMA
- IC-1154LLaMA-65B exhibits a selection bias where zero-shot accuracy varies substantially by answer choice, with 'a' at 71.58% and 'd' at 52.28%LLaMA
- IC-1155ViT-B/16 (ImageNet-21k) fine-tuned with VPT outperforms full fine-tuning on 16 of 19 VTAB-1k tasks, with the advantage concentrated in high-task-disparity and similar-distribution scenarios and narrowing as downstream data growsViT, Swin Transformer
- IC-1156The VPT advantage over FT for ViT-B/16 is not explained by overfitting resistance or additional optimization dimensions; the specific feature-preservation mechanism of VPT is the key factorViT
- IC-1157GPT-4 and other LMs show a large gap between rule induction and rule application, with task accuracy dropping to near zero on MiniScan when the LM itself applies its own proposed rulesGPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1158GPT-4 and other LMs are brittle to noisy exemplars and unfamiliar output representations, with performance degrading sharply even under minimal perturbationGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1159GPT-4 is a strong inductive hypothesis proposer, achieving high accuracy on inductive reasoning benchmarks when its generated rules are applied by a symbolic interpreterGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base
- IC-116In LLaVA-7B, multi-head attention modules drive hallucination more than MLP modules, and targeted intervention on specific hallucination heads reduces the hallucination rate by up to 1.7xLLaVA-1.5 / LLaVA-v1.5, MiniGPT-4
- IC-1165ProtoPFormer achieves 42.2% attribute identification accuracy on CUB-200-2011 in a 7-rater human evaluationProtoPFormer
- IC-1166MPLUG-Owl's VQA accuracy on VQA-X increases from 68.30% to 74.48% when prompted with progressively higher-quality rationales generated by RAPPERmPLUG-Owl
- IC-1167GPT-4 Code Interpreter's mathematical reasoning accuracy is positively correlated with code usage frequency, with its iterative code generation and self-debugging mechanism as the primary driver of its 69.69% zero-shot MATH accuracyGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1168For GPT-4 Code Interpreter, code-based self-verification improves MATH accuracy to 73.54% while natural language self-verification slightly degrades it to 69.29% relative to the 69.69% baseGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1169CodeLlama-7B and CodeLlama-34B show improvement with CSV prompting on GSM8K and MATH but at much lower absolute accuracy than GPT-4 Code InterpreterCodeLlama-13B, CodeLlama-34B
- IC-117Hallucination heads in LLaVA-7B and MiniGPT-4 are concentrated in the middle and deeper layers of the transformerLLaVA-1.5 / LLaVA-v1.5, MiniGPT-4
- IC-1170GPT-3.5, Llama2, PaLM2, and GPT-4 are susceptible to a CoT-prompting backdoor attack (BadChain) on complex reasoning tasks, with stronger reasoning models showing higher attack success ratesGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, PaLM 2, Llama 2 / Llama 2 base
- IC-1171Pre-trained ViT, MAE, and ResNet50 (supervised and MoCo v2) place visually similar but semantically distinct ImageNet classes (mop, broom, puck, crutch) in close proximity in their feature spaceViT, MAE, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-1175ChatGPT can be prompted to generate misinformation with near-perfect success for implicit methods but is largely resistant to explicit misinformation requestsChatGPT
- IC-1176ChatGPT-generated misinformation is harder for humans to detect than human-written misinformation with the same semanticsChatGPT
- IC-1177LLM-generated misinformation is harder for LLM detectors to detect than human-written misinformation with the same semanticsChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base, Vicuna
- IC-118Hallucination heads in LLaVA-7B and MiniGPT-4 allocate 4.75x more attention to text tokens than image tokens, and this pattern is inherited from the base language modelLLaVA-1.5 / LLaVA-v1.5, Vicuna, MiniGPT-4, Llama 2 / Llama 2 base
- IC-1188GPT-3 procedural planning performance is scale-dependent, with Curie (6.7B) scoring 3.75 and Davinci (175B) scoring 4.90 overall quality in few-shot settings, while GPT-4 achieves 4.81 overall and 5.00 orderGPT-3 / GPT base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-119The number of salient hallucination heads decreases as model size increases within the LLaVA familyLLaVA-1.5 / LLaVA-v1.5, LLaVA-NeXT / LLaVA 1.6
- IC-1193Vicuna and Alpaca achieve 0% pass rate on all ToolBench tool-use instructions, while GPT-4 and ChatGPT reach 71.1% and 64.8% with DFSDT, revealing a wide capability gap in tool use among released LLMsVicuna, Alpaca, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, ChatGPT, GPT-3 / GPT base
- IC-1194GPT-3.5-turbo, GPT-4, and GPT-3.5-turbo-0613 exhibit 50-58% inconsistency between their ratings and rankings feedback on the same response pairsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1195Model substitution adversarial attack reduces TreeRing AUROC to 0.14 at ε=2/255 and StegaStamp AUROC to 0.492 at ε=12/255TreeRing, StegaStamp
- IC-1196Blending a watermarked noise image with a clean image causes watermark detectors to falsely flag clean images as watermarkedRivaGAN, TreeRing
- IC-1199GPT-2 next-token distributions contain correctable tail errors from the softmax bottleneck that degrade generation quality under low-entropy sampling, with basis-aware threshold sampling improving MAUVE across all four sizesGPT-2
- IC-120YOLO-World can replace SAM as a 2D crop generator in OpenMask3D's pipeline with nearly equivalent mAP but 1.76x faster inferenceYOLO-World, SAM, YOLOv8, RT-DETR
- IC-1200GPT-2-XL's untruncated next-token log-probability matrix has rank saturating at its hidden dimensionality of 1600, while truncation sampling produces post-truncation distributions whose estimated rank grows far beyond 1600GPT-2
- IC-1201CLIP-ViT-L/14 image features support 200-way zero-shot EEG-based object recognition better than ViT-B/16 or ResNet-50 features when used as a frozen encoder in a contrastive learning frameworkCLIP / CLIP-ViT (LC), ViT, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, EVA-CLIP
- IC-1202Stable Diffusion 1.5 and 2.1 exhibit a cropping failure mode where synthesized objects are cut off at image boundariesStable Diffusion
- IC-1203Stable Diffusion 1.5 and 2.1 receive low user-preference win rates (7.91% and 6.71%) in a four-way comparison against SDXLStable Diffusion
- IC-1204SD-VAE 1.x and SD-VAE 2.x achieve lower reconstruction quality than the new SDXL-VAE on COCO 2017Stable Diffusion
- IC-1205GPT-3 models (ada, curie, davinci) achieve near-zero accuracy on zero-shot arithmetic tasks but learn them rapidly with 1000 fine-tuning samplesGPT-3 / GPT base
- IC-1206GPT-2-XL, GPT-J, Falcon-7B, Llama-2-7B, and Llama-2-13B are vulnerable to backdoor injection via lightweight parameter editing with only 15 samples, achieving near-100% attack success rate while preserving clean performanceGPT-2, GPT-J, Falcon, Llama 2 / Llama 2 base
- IC-1207For GPT-2-XL, backdoor injection via parameter editing is most effective on intermediate layers (15-35) and notably less effective on the first 10 and last 5 layersGPT-2
- IC-1208Larger LLMs (Llama-2-13B) require more data samples for successful backdoor injection via parameter editing compared to smaller models (GPT-2-XL 1.5B)GPT-2, Llama 2 / Llama 2 base
- IC-121Gemini 1.0 Pro, when prompted as a zero-shot chain-of-thought judge with majority voting, underperforms fine-tuned smaller Gemma models as verifiers on GSM8KGemini 1.0 Pro
- IC-1217LLaMA 65B's token-probability readout fails to capture human decision-making, producing near-chance NLL and no human-like exploration behaviorLLaMA
- IC-1218GPT-4 achieves 59.72% accuracy on choices13k and 80.3% on the horizon task when modeling human decisionsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1219Frozen GPT-2 XL contains pre-trained attention heads that implement the nearest-neighbor algorithmGPT-2
- IC-122Concept representations in Llama-2-7B, Gemma-7B, and Llama-2-13B become more consistent in deeper layersLlama 2 / Llama 2 base, Gemma
- IC-1220GPT-4, GPT-3.5-turbo, and Llama-2-70B can implement learning algorithms in-context on novel boolean functions, competing with nearest-neighbor baselinesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base
- IC-1221LLM performance on in-context boolean function learning is scale-dependent, with GPT-2 failing and Llama-2 models improving gradually with sizeGPT-2, Llama 2 / Llama 2 base
- IC-1225GPT-4 generates realistic dynamic scene layouts from text prompts with only 3 in-context examples, achieving 98% average accuracy across 5 spatiotemporal tasks, with physics knowledge (gravity, elasticity, perspective) generalising to unseen objects from its weightsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1226SAM's ViT-B encoder achieves 54.2% ImageNet-1k linear probing accuracy versus 67.7% for MAE's ViT-B, indicating its segmentation pretraining impairs high-level semantic representationSAM, MAE
- IC-1227SAM's segmentation pretraining shifts attention heads toward local focus in deeper layers, unlike its MAE initialization which retains global attention throughoutSAM, MAE
- IC-123Llama-2-7B, Gemma-7B, and Llama-2-13B organize 16 concepts into hierarchical clusters in their representation space that reflect real-world category structureLlama 2 / Llama 2 base, Gemma
- IC-1231On APPS, released code generation models span pass@1 from 0.20 (GPT-3 175B) to 6.20 (CodeRL), with value-based and policy-based RL methods outperforming supervised baselinesCodeRL, GPT-2, GPT-3 / GPT base, GPT-Neo, GPT-J, Codex, AlphaCode, PPOCoder
- IC-1232GPT-2 XL and GPT-J exhibit knowledge conflict when subjected to reverse and composite knowledge edits, with ROME and MEMIT showing near-total failure on reverse editsGPT-2, GPT-J
- IC-1233GPT-2 XL and GPT-J exhibit irreversible knowledge distortion after round-editing, with the effect being more severe when the edit target is semantically distant from the true labelsGPT-2, GPT-J
- IC-1236Among 7B LLMs, Llama-2-7b achieves the best zero-shot COMET scores in both translation directions, while MPT-7b leads in BLEU for en-to-xxLlama 2 / Llama 2 base, MPT, OPT, BLOOM, Falcon, Llama-1-7B, GPT-3.5 / ChatGPT-3.5
- IC-1237Llama-2-7b's pre-existing translation knowledge is diluted by large amounts of parallel data, causing COMET to decline after 100k examplesLlama 2 / Llama 2 base, MPT
- IC-1238Llama-2-13b produces off-target non-translation outputs in zero-shot English-to-foreign-language translationLlama 2 / Llama 2 base
- IC-1239Llama-2-7b and Llama-2-13b achieve top zero-shot cross-lingual performance among 7B models on XNLI, XStoryCloze, and XWinogradLlama 2 / Llama 2 base, Xlm-R, XGLM, BLOOM, MPT
- IC-124Direct comparison of CLIP image embeddings with CLAP audio embeddings achieves near-chance retrieval, while logsumexp bridging through the shared language modality recovers 62% recall@10 on AudioSetCLIP / CLIP-ViT (LC), CLAP
- IC-1240GPT-4-0314 achieves 85% zero-shot accuracy on situational-awareness questions about its own architecture and trainingGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1241GPT-4 (14 March 2023) achieves 100% zero-shot accuracy at classifying whether news articles could be part of its pre-training dataGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1242GPT-4 wrote a working script that called an instance of itself on its API as part of a plan to gain internet accessGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1243GPT-4 and GPT-4 Turbo achieve near-zero scores on GAIA level 3 and single-digit to low-double-digit scores on levels 1-2, compared to 87-94% for human annotatorsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1244GPT-4's non-zero scores on GAIA web browsing questions are largely due to memorization of intermediate information from training data rather than actual web browsingGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1249TD-MPC exhibits training instability and performance degradation when the planning horizon is set to 20 time stepsTD-MPC
- IC-125LanguageBind's direct evaluation achieves 70% recall@10 on AudioSet, empirically validating that the inner product between unpaired modality representations recovers the correct probability ratioLanguageBind
- IC-1250GPT-2-medium shares 78% of its top attention heads between the IOI circuit and the colored objects circuitGPT-2
- IC-1251Intervening on four attention heads in GPT-2-medium boosts colored objects accuracy from 49.6% to 93.7% by making the circuit behave like the IOI circuitGPT-2
- IC-1252Circuit overlap between IOI and colored objects in GPT-2 decreases as model scale increases from medium to xlGPT-2
- IC-1256MPT-7B-Chat produces non-committal responses rather than proper refusals on unsafe instructionsMPT
- IC-1257Guanaco acknowledges the illegality of requested actions but still provides the harmful informationGuanaco
- IC-126CLIP and CLAP language representations are statistically indistinguishable from a uniform distribution on the hypersphereCLIP / CLIP-ViT (LC), CLAP
- IC-1261LLaMA-2 attention to constraint tokens correlates with factual correctness, and a linear probe on these attention weights predicts factual errors comparably to model confidenceLlama 2 / Llama 2 base
- IC-1262LLaMA-2 factual query accuracy improves with entity popularity and decreases with query constrainedness, with larger models showing better performance on less popular and more constrained queriesLlama 2 / Llama 2 base
- IC-1263LLaMA-2 7B and 13B attention signal for predicting factual errors is available by approximately 50% of layers, enabling early stopping without performance degradation, while LLaMA-2 70B shows a slight performance dropLlama 2 / Llama 2 base
- IC-1264LLMs are overconfident when verbalizing confidence, with values concentrated in 80–100% and multiples of 5, yielding high ECE across all five tested modelsGPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Vicuna, Llama 2 / Llama 2 base
- IC-1265Calibration and failure prediction improve as model capability scales from GPT-3 to GPT-4, but remain far from idealGPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Vicuna, Llama 2 / Llama 2 base
- IC-1266For GPT-3, white-box token-probability methods outperform black-box verbalized confidence in uncertainty estimation, but the gap is narrow (0.522–0.605 AUROC) and both remain near randomGPT-3 / GPT base
- IC-1267LLMs' alignment with human privacy judgments drops sharply as contextual complexity increases from tier 1 to tier 3GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, ChatGPT, InstructGPT, Mixtral, Llama 2 / Llama 2 base
- IC-1268LLMs leak private information in theory-of-mind scenarios even when explicitly instructed to preserve privacyGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, ChatGPT, InstructGPT, Mixtral, Llama 2 / Llama 2 base
- IC-1269LLMs leak secrets to inappropriate recipients in meeting summarization and action-item generation tasksGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, ChatGPT, InstructGPT, Mixtral, Llama 2 / Llama 2 base
- IC-127ImageBind's direct evaluation closely matches logsumexp for both vision-language and audio-language alignment, validating the law for ImageBindImageBind
- IC-1270Chain-of-thought prompting does not mitigate privacy leakage in GPT-4 or ChatGPTGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, ChatGPT
- IC-128Released VLMs exhibit sycophancy, agreeing with incorrect user opinions while ignoring visual evidence, with LLaVA-1.5 showing the highest rate (94.6%) and InternLM-XComposer2-VL-1.8B the lowest (28.8%)BLIP-2, InstructBLIP, LLaVA-1.5 / LLaVA-v1.5, mPLUG-Owl2, InternVL-1.5, InternLM-XComposer2-VL, Gemini, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1283CLIP's InfoNCE training objective is mathematically equivalent to performing generalized spectral clustering on the bipartite image-text pair graphCLIP / CLIP-ViT (LC)
- IC-1284Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on as few as 10 explicitly harmful examples removes their safety alignment, raising harmfulness rates to 80-92%GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base
- IC-1285Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on 10 implicitly harmful identity-shifting examples (containing no toxic content) jailbreaks their safety alignmentGPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base
- IC-1286Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on benign utility-oriented datasets (Alpaca, Dolly, LLaVA-Instruct) degrades their safety alignment without any malicious intentGPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base
- IC-1287A backdoor can be implanted in GPT-3.5 Turbo via fine-tuning that is undetectable by standard safety auditing: the model appears safe on plain prompts but fulfills harmful instructions when a 3-word trigger is appendedGPT-3.5 / ChatGPT-3.5
- IC-1288GPT-4, ChatGPT, and GPT-4V fail to close the human-machine gap on Bongard-OpenWorld, with InstructBLIP captions differentially degrading ChatGPT while improving GPT-4ChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, BLIP-2, InstructBLIP
- IC-1289OpenFlamingo and Otter achieve near-chance accuracy on Bongard-OpenWorld, indicating inability to perform multi-image reasoningOpenFlamingo, Otter
- IC-129Sycophancy in VLMs increases with model size: InternVL-1.5-26B (95.8/89.6/86.5%) is more sycophantic than InternVL-1.5-2B (75.6/66.8/98.1%), and InternLM-XComposer2-VL-7B (36.7/28.0/50.7%) more than the 1.8B variant (33.3/20.2/33.0%)InternVL-1.5, InternLM-XComposer2-VL
- IC-1290CLIP, DINO, and DINOv2 as zero-shot natural baselines score below the 50% chance level on Bongard-OpenWorld due to adversarial query selectionCLIP / CLIP-ViT (LC), DINO, DINOv2
- IC-1293The l1 path-norm of PyTorch's pretrained ResNets is approximately 30 orders of magnitude too large for the path-norm generalization bound to be informative on ImageNet-1kResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-130Amplifying visual-token attention in high layers (16-32) of released VLMs reduces sycophancy while preserving VQA accuracy, indicating that insufficient high-layer visual attention is a key cause of sycophancyLLaVA-1.5 / LLaVA-v1.5, BLIP-2, InstructBLIP
- IC-1307Released LLMs (CodeLlama 7B/13B/34B, GPT-3.5, GPT-4) achieve limited code-optimization speedups with standard prompting, with the best baseline (GPT-3.5 CoT) reaching only 1.60x versus the 3.66x human referenceCode Llama, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1308Dynamic retrieval-based few-shot prompting substantially improves released LLMs' code optimization, with GPT-4-0613 reaching 76.07% optimization rate and 3.93x speedup (best@8), exceeding the 3.66x human referenceCode Llama, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1309GPT-4-0613 exhibits reduced output diversity relative to GPT-3.5: it outperforms on best@1 but underperforms on best@8 under CoT promptingGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-131GPT-4 Turbo, GPT-3.5 Turbo, Llama3-8B, Qwen-7B, and iFlytekSpark-13B over-rely on the strong reminder 'the answer is' in prompts as a shortcut, with accuracy dropping sharply when the cue is a random answer rather than the ground truthGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 3, Qwen, iFlytekSpark-13B
- IC-1310Chain-of-thought prompting provides notable code-optimization gains only for larger models (CodeLlama 34B, GPT-3.5, GPT-4) but not for CodeLlama 7B or 13B, consistent with an emergent capabilityCode Llama, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1317Llama-2 and Pythia models contain linear representations of space and time that improve with depth and model scaleLlama 2 / Llama 2 base, Pythia
- IC-1318Individual space and time neurons in Llama-2-7B causally contribute to spatial and temporal predictionsLlama 2 / Llama 2 base
- IC-1319Larger LLaMA and LLaMA2 models show better calibration on phrase-level tasks but not consistently on sentence- and paragraph-level tasksLLaMA, Llama 2 / Llama 2 base
- IC-132GPT-4 Turbo, GPT-3.5 Turbo, Llama3-8B, Qwen-7B, and iFlytekSpark-13B trust authority roles (teacher/judge) more than peer roles (classmate/lawyer) when the cue information is the correct answerGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 3, Qwen, iFlytekSpark-13B
- IC-1320GPT-2 XL (1.5B) exhibits better calibration than larger models from the LLaMA, LLaMA2, and GPT-J families despite having fewer parametersGPT-2, GPT-J, LLaMA, Llama 2 / Llama 2 base, Vicuna
- IC-1321Vicuna-13B, instruction-tuned from LLaMA-13B on user conversations, exhibits worse calibration than its base model LLaMA-13BVicuna, LLaMA
- IC-1327GPT-4, GPT-3.5, Llama2, and Vicuna models underperform human annotators on multistep soft reasoning in natural language narratives, with smaller models scoring near random chanceGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base, Vicuna
- IC-1328GPT-4's performance on MUSR depends on the type of neurosymbolic scaffolding: program-aided decomposition helps on structured optimization but symbolic belief tracking fails on natural language theory-of-mindGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-133GPT-4, Claude 3, and Gemini 1.0 Pro do not exhibit detectable watermarks from the red-green, fixed-sampling, or cache-augmented families under black-box statistical testsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3, Gemini 1.0 Pro
- IC-1330All 20 evaluated LLMs improve in multi-turn task-solving with additional tool-use turns and GPT-4-simulated language feedbackGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude Instant 1, Chat-bison-001, Llama 2 / Llama 2 base, Vicuna, CodeLlama-13B, CodeLlama-34B, Lemur-v1-70B / Lemur-70B-Chat-V1
- IC-1331SIFT and RLHF variants of CodeLlama and Llama-2 perform worse than their base counterparts in multi-turn interactionCodeLlama-13B, CodeLlama-34B, Llama 2 / Llama 2 base, Vicuna, Lemur-v1-70B / Lemur-70B-Chat-V1
- IC-1332Vicuna-v1.5 and CodeLlama-34b-instruct produce format-breaking artifacts (escaped underscores, [python] tags) in 30-100% of code instances due to training data contaminationVicuna, CodeLlama-34B
- IC-1336MobileNetV2 (PyTorch pre-trained on ImageNet) exhibits a failure mode under global unstructured L1 pruning at the Pareto-optimal point, with its high kurtosis of kurtoses (64.40) causing very-low-magnitude layers to be entirely pruned and disconnect the networkMobileNetV2, VGG / VGG13, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-1337InstructPix2Pix is more effective at editing color than at preserving category in visual concept editingInstructPix2Pix
- IC-1338Chinchilla 70B and Llama 2 7B, trained primarily on text, compress ImageNet patches and Librispeech audio better than domain-specific compressors PNG and FLACChinchilla, Llama 2 / Llama 2 base
- IC-1339Chinchilla 1B's compression rate improves with increasing sequence length across text, image, and audio, demonstrating in-context learning without gradient updatesChinchilla
- IC-134Mistral Large V2 as a judge on SummEval coherence systematically avoids extreme ratings (1 and 5), while human annotators assign over 24% of items a median rating of 5Mistral Large V2
- IC-1340Chinchilla 70B produces coherent autoregressive continuations of text, image, and audio data when used as a compressor, outperforming gzip in sample qualityChinchilla
- IC-1341DRUM's standard datalog rule extraction is fundamentally incomplete because its predictions depend on counting distinct rule-body matchesDRUM
- IC-1342Assigning socio-demographic personas to LLMs causes significant reasoning performance degradation across all four models studied, manifesting as both explicit abstentions and implicit reasoning errorsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base
- IC-1343Task-agnostic de-biasing prompts are ineffective at reducing persona-induced reasoning bias in ChatGPT-3.5, while task-dependent expertise prompts are effective but lack generalizabilityGPT-3.5 / ChatGPT-3.5
- IC-135Linear relational embeddings for factual relations form in OLMo-7B, OLMo-1B, and GPT-J when subject-object co-occurrence frequency exceeds model-specific thresholds, with r=0.82 correlation between log co-occurrence and causality across all pretraining stagesOLMo / OLMo base, GPT-J
- IC-1355Instruction-tuned VLMs fail to follow multiple-choice format in reasoning questions, with InstructBLIP frequently returning blank responsesInstructBLIP, LLaVA, LLaMA-Adapter v2, mPLUG-Owl, Otter
- IC-1356GPT-3.5 turbo achieves over 90% agreement with human judgments when evaluating VLM responses on open-set questionsGPT-3.5 / ChatGPT-3.5
- IC-136LRE quality metrics from OLMo-7B predict pretraining term frequencies in GPT-J (trained on different data) with approximately 70% within-magnitude accuracy for object frequencies, outperforming log-probability-only features by about 30%OLMo / OLMo base, GPT-J
- IC-1361GPT-4 and other state-of-the-art LLMs achieve near-human accuracy in inferring personal attributes from unstructured textGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base, Claude Instant 1, PaLM 2
- IC-1362State-of-the-art text anonymization is insufficient to prevent GPT-4 from inferring personal attributesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1363Current model alignment does not filter privacy-invasive prompts across major LLM providersLlama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, PaLM 2
- IC-1364GPT-4 can extract personal information from users through adversarial chatbot conversationsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1369Successor heads that increment ordinal-sequence tokens exist in Pythia, GPT-2, and Llama-2 models from 31M to 12B parametersPythia, GPT-2, Llama 2 / Llama 2 base
- IC-137Pre-trained ResNet34 and ViT-B features on CIFAR-100 exhibit a block-diagonal class-correlation structure, with ViT-B showing higher intra-class correlation (0.35) than ResNet34 (0.25)ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, ViT
- IC-1370MLP0 representations of ordinal-sequence tokens in Pythia-1.4b contain linearly decodable mod-10 features that are causally important for incrementationPythia, GPT-2
- IC-1371The Pythia-1.4b successor head l12h0 exhibits interpretable polysemanticity, performing successorship, acronym prediction, copying, and greater-than behaviors on natural language dataPythia
- IC-1372Successor heads in Pythia-1.4b exhibit a greater-than bias: the OV circuit assigns systematically higher logits to tokens with greater ordinal values than the input, impairing decrementationPythia
- IC-1373All five evaluated LLMs show a strong positional bias in constrained text generation, with first-position constraints nearly always satisfied but last- and arbitrary-position constraints causing major performance dropsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, PaLM 2, Vicuna
- IC-1374Counting difficulty in constrained generation increases with text level and constraint strictness, with exact sentence-level character counts being the hardest condition for all five modelsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, PaLM 2, Vicuna
- IC-1375GPT-4's constraint satisfaction improves by approximately 20% after one round of automated feedback but plateaus at 66% even after three additional roundsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-138Trojan backdoored Llama-2-7B models and Vicuna-7B-v1.5 exhibit the probe concatenate effect, where concatenating a triggered or jailbroken sample with a harmful probe significantly shifts the model's output distribution away from refusalVicuna
- IC-1380Pythia and OPT small models exhibit non-trivial performance on BigBench tasks that is invisible under beam search but revealed by extensive random samplingPythia, OPT
- IC-1386Fact recall in OPT and LLaMA models degrades by more than 5% relative accuracy when more than 30% of weights are pruned, and similarly when moving from the 30B to the 13B dense modelOPT, LLaMA, Pythia
- IC-1387In-context learning capabilities in OPT and LLaMA models remain within 5% of dense-model accuracy even at 60-70% sparsity, and show less than 2% difference between the 30B and 1.3B dense OPT modelsOPT, LLaMA, Pythia
- IC-1388In LLaMA-13B, feed-forward layers are more critical than attention layers for fact recall, while both are equally important for in-context learningLLaMA
- IC-1389Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMACLIPBERT, UniVL, VideoCLIP, CLIP4Clip, VioLET, X-CLIP, UniPerceiver, Merlot Reserve, VindLU, InternVideo, mPLUG-2, Otter, Video-LLaMA, CLIP / CLIP-ViT (LC), BLIP-2, GPT-2, OPT
- IC-1393D Gaussian Splatting and its variants (Scaffold-GS, Mip-Splatting) are vulnerable to computation cost attacks via data poisoning, with peak GPU memory increasing up to 21.93x and training time up to 4.97x under unconstrained perturbation3D Gaussian Splatting, Scaffold-GS, Mip-Splatting
- IC-1390Proficiency tests reveal that a substantial portion of correct main-test predictions by VidLMs and ILMs are spurious rather than reflecting robust understandingCLIPBERT, UniVL, VideoCLIP, CLIP4Clip, VioLET, X-CLIP, UniPerceiver, Merlot Reserve, VindLU, InternVideo, mPLUG-2, Otter, Video-LLaMA, CLIP / CLIP-ViT (LC), BLIP-2, GPT-2, OPT
- IC-1391SLD concept removal variants and SD with negative prompts are bypassable by Ring-a-Bell adversarial prompts, increasing attack success rate from single digits to 90-100% for nuditySLD-max, SLD-strong, SLD-medium, Stable Diffusion
- IC-1392Title reproduction shows no contamination signal while tag reproduction shows a negative association with GitHub presence and a moderating difficulty effect for GPT-4 and GPT-3.5-turboGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-1393Most mainstream LLMs generate value-violating content at high rates (APV 65-80%) across 2,397 morally ambiguous prompts, indicating substantial ethical misalignmentLlama 2 / Llama 2 base, LLaMA, ChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Falcon, Vicuna, Guanaco, Baichuan, Baichuan 2, Baichuan2-13B, ChatGLM-6B / ChatGLM-6b-2, GPT-3 / GPT base
- IC-1394ChatGPT demonstrates better ethical value conformity than GPT-4 across multiple prompt generation sourcesChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1395ChatGPT's ethical violation rate decreases from 70.07 to 57.58 APV when given targeted in-context value instructions generated by VILMO, outperforming baseline alignment methodsChatGPT, Llama 2 / Llama 2 base, GPT-3 / GPT base
- IC-1399Invariant GNNs (l=0) consistently fail to distinguish k-hop identical but globally distinct geometric graphs on the k-chain task, regardless of model depthSchNet, DimeNet++, SphereNet, CoMEt, MACE, GVP, EGNN, CloFNet, ESCN, EquiformerV2
- IC-1400When steerable feature dimension is held constant, increasing the type-l of steerable features does not improve performance of ESCN or EquiformerV2 on IS2RE and S2EF molecular property predictionESCN, EquiformerV2
- IC-1401GPT-4 serves as a proxy for human judgment on SOTOPIA-EVAL, with strong correlations on goal, financial, and relationship dimensions for model outputsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1402Llama-2-70B-Chat underperforms GPT-3.5 across all SOTOPIA dimensions in interactive social scenarios, diverging from static benchmark rankingsLlama 2 / Llama 2 base, GPT-3.5 / ChatGPT-3.5
- IC-1403All four evaluated LLMs produce negative scores on social rules and secret-keeping dimensions in SOTOPIA interactionsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base, MPT
- IC-1404On SOTOPIA-HARD, GPT-4 achieves significantly lower goal completion than humans and exhibits non-strategic negotiation and excessive compromise behaviorsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1405OpenFlamingo and Idefics models hallucinate objects not present in images, and increasing ICL shots beyond 4 amplifies hallucinationsOpenFlamingo, Idefics
- IC-1406OpenFlamingo and Idefics models rarely abstain from answering unanswerable questions, but ICL significantly improves abstention F1OpenFlamingo, Idefics
- IC-1407OpenFlamingo and Idefics models perform near random chance on compositional image-text matching, and ICL has almost no effect on atomic foilsOpenFlamingo, Idefics
- IC-1408OpenFlamingo and Idefics models generate low-quality explanations in zero-shot, but ICL and model scale significantly improve explanation CIDErOpenFlamingo, Idefics
- IC-1409FF blocks in BERT and GPT-2 modify token-to-token contextualization, with the effect concentrated in specific layers and targeting specific linguistic compositions rather than simple word co-occurrenceBERT, MultiBERTs, RoBERTa / RoBERTa-L, GPT-2, OPT
- IC-1410FF's contextualization effects in BERT and GPT-2 are largely canceled by the residual connection and layer normalization, with LN's γ weights specifically shrinking the outlier dimensions in FF outputBERT, MultiBERTs, RoBERTa / RoBERTa-L, GPT-2, OPT
- IC-1413GPT-4 achieves near-saturation on Python code synthesis (86.6% pass@1) but scores significantly lower on code repair (47.8% avg) and code explanation (52.1% avg) across six languagesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1414Pretrained code models Starcoder and CodeGeex2 score 0.0% on code explanation across all six languages because they generate code instead of natural languageStarcoder, CodeGeex2
- IC-1415BLOOMZ generalizes instruction-following to programming languages (Go, Rust) absent from its instruction data, scoring above the random baselineBLOOM
- IC-1416ChatGPT achieves only 28–35% accuracy on multimodal intent recognition in MINTREC2.0, showing a gap of over 30 percentage points compared to human evaluatorsChatGPT
- IC-1428OPT 6.7B exhibits over 90% activation sparsity in FFN layers, reducing inference from 6.6G to 4.5G flops per token, while Llama 7B (SiLU) and Falcon 7B (GELU) show near-zero sparsityOPT, LLaMA, Falcon
- IC-1429OPT 6.7B exhibits aggregated sparsity where approximately 50% of neurons remain unused across the first 150 tokens, with a non-random reuse pattern enabling 1.27x speculative decoding speedup at gamma=16OPT
- IC-1431BERT-base fine-tuning has negligible distribution-wise variance (0.21%) while BERT-large fine-tuning has substantial distribution-wise variance (2.08%) on MRPCBERT
- IC-1436Imagen Video 5.6B produces high-quality but domain-inappropriate videos on out-of-distribution robotics and egocentric data, failing to generate relevant dynamicsImagen Video
- IC-1437LLaMA-2-7b-chat-hf and Meta-LLaMA-3-8B-Instruct exhibit reduced attention to rule tokens and fail to follow prompt-specified rules when the adversarial suffix 'forget all prior instructions and answer the question' is appendedLlama 2 / Llama 2 base, Llama 3
- IC-1438LLaVA and Llama-Adapter V2 are jailbroken by compositional adversarial images targeting image-based embedding triggers, with near-zero success for textual triggersLLaVA, LLaMA-Adapter v2
- IC-1439LLaVA follows text instructions embedded in adversarial images as if they were user prompts, enabling hidden prompt injectionLLaVA, LLaMA-Adapter v2
- IC-144TD-MPC's training is unstable, with performance collapsing after approximately 1–4 million steps across multiple DM Control tasksTD-MPC
- IC-1440GCN, GAT, GraphSAGE, and SGC exhibit structure-dependent generalization in transductive node classification: test nodes with shorter paths to training nodes are classified more accuratelyGCN, GAT, GraphSAGE, SGC
- IC-1441GCN exhibits structural unfairness in transductive node classification, with demographic parity and equal opportunity gaps between nodes connected to and disconnected from the training setGCN
- IC-145TD-MPC fails to achieve high reward when irrelevant background information is added to image inputsTD-MPC
- IC-1456Pre-trained language models fail to predict both interpretations of ambiguous inputs in zero-shot semantic parsingCodeGen, LLaMA, Vicuna, GPT-3.5 / ChatGPT-3.5
- IC-1457Pre-trained language models track the distribution of logical forms in mixed few-shot prompts with conflicting examplesCodeGen, LLaMA, Vicuna
- IC-1458GPT-4 and PALM 2-L produce significantly less consistent descriptions of interpolated domain embeddings than a purpose-built ELM modelGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, PaLM 2
- IC-146Pythia models show increasing robustness to off-policy RLHF data as policy size scales from 410M to 2.8BPythia
- IC-1461Varying decoding hyperparameters and removing the system prompt breaks the safety alignment of 9 out of 11 open-source LLMs, raising attack success rate from 0% to over 95%Vicuna, MPT, Falcon, Llama 2 / Llama 2 base
- IC-1462GPT-3.5-turbo is substantially more robust to the generation exploitation attack, with attack success rate of only 7% compared to over 95% for open-source modelsGPT-3.5 / ChatGPT-3.5
- IC-1463ChatGPT achieves 34.9% average F1 on zero-shot NER across 43 datasets spanning 9 domainsChatGPT
- IC-1464Vicuna-7B and Vicuna-13B achieve only 14.2% and 18.0% average F1 on zero-shot NER, trailing ChatGPT by over 20 pointsVicuna
- IC-1465InstructUIE-11B achieves 81.16% average F1 on 20 in-domain NER datasets and 49.4% on out-of-domain evaluationInstructUIE
- IC-1467DECAF achieves PVE 9.65 and f-score 89.6 on the DECAF validation set with a runtime of 19.59 seconds per imageDECAF
- IC-1468DECAF's optimization-based fitting degrades under significant self-occlusion where the hand covers more than half the faceDECAF
- IC-1469Progress on standard ImageNet generalization benchmarks is 2.5x faster than progress on crowdsourced global data (DollarStreet, GEODE) across 98 vision modelsCLIP / CLIP-ViT (LC), ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, DINOv2, ViT, ConvNeXt, RegNet, MobileNetV3, VGG / VGG13, HRNet, FLAVA, EVA-CLIP, MLP-Mixer, EdgeNeXt, ReXNet
- IC-147CLIP's OOD performance on rendition domains is largely an artifact of domain contamination in its web-scale training dataCLIP / CLIP-ViT (LC)
- IC-1470Geographic disparities (Europe-Africa accuracy gap) are large across all 98 models and have more than tripled between least and best performing models on DollarStreetCLIP / CLIP-ViT (LC), ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, DINOv2, ViT, ConvNeXt, RegNet, MobileNetV3, VGG / VGG13, HRNet, FLAVA, EVA-CLIP, MLP-Mixer, EdgeNeXt, ReXNet
- IC-1471Common robustness interventions (AugMix, CutMix, Deep AugMix, texture debiasing, antialiasing) and scaling of data or model size do not resolve geographic disparities in released vision modelsResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, CLIP / CLIP-ViT (LC)
- IC-1472DINOv2 (86M parameters) achieves the smallest GEODE geographic disparity (2.46% Europe-Africa gap) among all 98 models in the testbedDINOv2
- IC-1477TAPe performs poorly compared to a simple CNN for protein fitness predictionTAPe
- IC-148Language models represent semantically equivalent inputs from different data types (languages, code, images, audio) close together in intermediate layers, with the shared space scaffolded by the model's dominant languageLlama 2 / Llama 2 base, Llama 3, Baichuan 2, BLOOM, LLaVA, Chameleon, SALMONN
- IC-1484GP-UNIT's FID degrades under noisy inputs in reference-guided mode but paradoxically improves in latent-guided modeGP-UNIT
- IC-1485Sketch Transformer's FID degrades from 31.49 to 404.01 under Gaussian noise at the highest tested intensitySketch Transformer
- IC-1486HiFaceGAN's FID degrades from 34.83 to 320.41 under Gaussian noise at the highest tested intensity for face super-resolutionHiFaceGAN
- IC-1487CycleGAN's FID degrades from 76.92 to 180.82 under Gaussian noise for horse-to-zebra translationCycleGAN
- IC-1489State-of-the-art foundation models (CLIP, GPT-3.5-turbo, and others) score well below elementary students on multimodal K-12 STEM questionsCLIP / CLIP-ViT (LC), GPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5, ViLBERT, 12-in-1, UNITER, ViRTex, UnifiedQA, GloVe
- IC-149Intervening in the shared representation space using the dominant language (English) predictably changes model outputs for other data types, demonstrating the space is causally used rather than a vestigial byproductLlama 3, Llama 2 / Llama 2 base, Chameleon, SALMONN
- IC-1490Zero-shot CLIP is overconfident on STEM questions, with softmax confidence loosely related to actual accuracyCLIP / CLIP-ViT (LC)
- IC-1491CLIP zero-shot performance on STEM saturates across model sizes, with only 3.6 points of variation from smallest to largest variantCLIP / CLIP-ViT (LC)
- IC-1496A single frozen transformer block from LLaMA-7B consistently improves performance across diverse visual tasks when appended to existing visual encodersLLaMA
- IC-1497LLaMA-7B's frozen transformer block amplifies informative visual tokens, producing feature activations with higher Miou against ground-truth segmentation masks than both the baseline ViT and the model's own attention scoresLLaMA
- IC-1498The benefit of frozen LLM transformer blocks for visual encoding is scale-dependent: OPT blocks below 1.3B parameters degrade ViT-s performance while blocks at 1.3B and above improve itOPT
- IC-1499Self-rationalization quality and task accuracy scale with model size across GPT-3, FLAN-T5, and LLaMA on five QA datasetsGPT-3 / GPT base, FLAN-T5, LLaMA
- IC-150Second-order effects of CLIP's MLP neurons are concentrated in late layers (8–10 of 12 in ViT-B/32)CLIP / CLIP-ViT (LC)
- IC-1508LLMs with in-context learning translate Kalamang-English at 44.7/45.8 CHRF, falling short of the human baseline of 51.6/57.0 CHRFGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, GPT-3 / GPT base, Llama 2 / Llama 2 base, LLaMA
- IC-1509Kalamang-English translation performance on MTOb increases with model size within the Llama and Llama 2 families, and GPT-4 outperforms Text-davinci-003LLaMA, Llama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3 / GPT base
- IC-151Each CLIP neuron's second-order effect is approximately a single linear direction in the joint text-image space, significant for fewer than 2% of imagesCLIP / CLIP-ViT (LC)
- IC-1510Without retrieved context, LLMs are unable to translate Kalamang, and among context types, retrieved parallel sentences are most beneficial, followed by word list entries, then grammar book passagesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, GPT-3 / GPT base, Llama 2 / Llama 2 base, LLaMA
- IC-1518Domain finetuning of LLaMA 2 7B, LLaMA 2 13B, and GPT-2 XL on PubMed causes topic and style priors to shift dramatically, accounting for the majority of the probability change, while factual knowledge learning contributes only a small fractionGPT-2, Llama 2 / Llama 2 base
- IC-1519Topic and style biases in LLaMA 2 7B are learned like simple features (rapidly, with minimal capacity, concentrated at the first few tokens, magnified by learning rate) while factual knowledge is learned like complex features (slowly, requiring significant capacity, uniformly across positions, unaffected by learning rate)Llama 2 / Llama 2 base
- IC-152CLIP's polysemantic neurons encode spurious correlations between unrelated concepts that can be exploited to generate adversarial misclassificationsCLIP / CLIP-ViT (LC)
- IC-1520OpenCLIP's per-sample zero-shot accuracy on ImageNet-based OOD benchmarks is strongly correlated with the perceptual similarity between that sample and its nearest neighbor in LAION-400MOpenCLIP
- IC-1523Five AI assistants (Claude-1.3, Claude-2.0, GPT-3.5-turbo, GPT-4, Llama-2-70B-Chat) consistently exhibit sycophancy across four varied free-form text-generation tasksClaude 1.3, Claude 2.0, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base
- IC-153ResNet50 relies on flower petals and green background features as shortcuts when classifying bee imagesResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-154CLIP ViT-L/14 text embeddings fail to capture fine-grained visual class similarities, ranking rottweiler and doberman at position 828 behind unrelated pairsCLIP / CLIP-ViT (LC)
- IC-1540CLS-token attention maps in pretrained ViT-t/16 exhibit high inter-layer correlation (cosine similarity up to 0.97) concentrated in layers 3–10, and MSA block outputs show high CKA in layers 2–8ViT
- IC-1543VGG19, ResNet50, ViT-Base, and DeiT-Base (ImageNet pretrained) achieve near-zero accuracy under query-based black-box attacks with 1000–10000 queriesVGG / VGG13, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, ViT, DeiT
- IC-1544The latent spaces of pretrained foundational models across vision and text are not related by a single class of geometric transformations; the optimal alignment depends on the specific model pair, architecture, and dataset.CLIP / CLIP-ViT (LC), ViT, RexNet 100, BERT-base-cased, BERT, ELECTRA-base-discriminator, RoBERTa / RoBERTa-L, ALBERT-base-v2, XLM-RoBERTa-base
- IC-1549All 28 evaluated LMs exhibit gender bias on non-stereotypical sentence pairs, with fairness scores between 9% and 41%Pythia, GPT-J, OPT, Llama 2 / Llama 2 base, MPT, OLMo / OLMo base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-155All 13 evaluated MLLMs perform at or near random guessing on MediConfusion, with confusion scores often exceeding 90%, indicating they cannot distinguish visually dissimilar radiology image pairsLLaVA, BLIP-2, InstructBLIP, DeepSeek-VL2, Molmo, LLaVA-Med, RadFM, Med-Flamingo, GPT-4o, O1 / OpenAI-o1-preview, Claude 3, Gemini 1.5 / Gemini Pro 1.5, Gemini
- IC-1550All evaluated LMs systematically prefer male pronoun completions in the non-stereotypical portions of Winobias and Winogender, with margins exceeding 40%Pythia, GPT-J, OPT, Llama 2 / Llama 2 base, MPT, OLMo / OLMo base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-1551No consistent relationship between model size and gender fairness scores is observed across six LM familiesPythia, OPT, Llama 2 / Llama 2 base, MPT, OLMo / OLMo base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-1552Deduplication of pretraining data does not consistently improve gender fairness in Pythia modelsPythia
- IC-1553GPT-J, GPT-2-XL, and Llama-13B decode approximately 48% of tested relations via a linear transformation on the subject representation, and this structure causally influences predictionsGPT-J, GPT-2, LLaMA
- IC-1554LRE faithfulness in GPT-J is concentrated in intermediate layers and drops sharply in later layers, consistent with a mode switch from relational encoding to next-token predictionGPT-J
- IC-1555GPT-J's internal representations contain correct factual knowledge even when the model outputs falsehoods under repetition or instruction distraction promptsGPT-J
- IC-1558Code LLaMA 13B maintains 99.4% passkey retrieval at 128k context despite perplexity rising from 2.37 to 2.54 between 98304 and 131072 tokensCodeLlama-13B
- IC-156Gemini models show substantially lower confusion scores than other MLLMs yet still perform at or near random guessing, suggesting their bottleneck is medical knowledge or reasoning rather than visual encodingGemini 1.5 / Gemini Pro 1.5, Gemini
- IC-157GPT-4o's MediConfusion performance is robust to prompt format while InstructBLIP is highly sensitive and LLaVA-Med fails completely on multiple-choice evaluationGPT-4o, InstructBLIP, LLaVA-Med
- IC-1576Base and aligned LLMs share 77.7% of top-1 token predictions, with distribution shifts concentrated in stylistic tokens rather than knowledge contentLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Vicuna
- IC-1577Base LLMs prompted with URiAL (3 restyled in-context examples + system prompt) match or surpass their SFT/RLHF-aligned counterparts on multi-aspect evaluationMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 2 / Llama 2 base, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1578XLM-R-XL without instruction tuning produces [pad] tokens and fails to complete instruction-following tasksXlm-R
- IC-1579ADM's noise prediction network exhibits exposure bias: during iterative sampling the l2-norm of its ε prediction is systematically larger than during training, and the sampling distribution variance exceeds the training variance with error accumulating toward the end of the chainADM
- IC-158Fine-tuning LLaVA-Med on MediConfusion training pairs cannot achieve 100% training accuracy, indicating the vision encoder's embeddings are fundamentally ambiguous for the confusing pairsLLaVA-Med
- IC-1580The official Llama 2-7B checkpoint fails to generate valid numerical responses for 3D-dependent molecular properties, with a valid answer rate of only 23% for SCF energyLlama 2 / Llama 2 base
- IC-1584LLaMA-7B and GPT-J-6B fail to interpret textual emphasis markers, with marked prompting degrading performance substantiallyLLaMA, GPT-J
- IC-1585LLaMA-7B and GPT-J-6B exhibit positional bias in instruction following: zero-shot performance varies significantly when the instruction is moved from after to before the contextLLaMA, GPT-J
- IC-1586In LLaMA-7B, steering all attention heads degrades JSON format accuracy below zero-shot, while steering a subset of 50-100 heads selected via multi-task profiling raises it to 96.64; performance varies dramatically across the 32 layers and individual headsLLaMA
- IC-159GPT-4o achieves 55.6% accuracy on creation, 74.8% on math, and 68.1% on code as a preference judge, and is outperformed by domain-specific 7B models on those tasksGPT-4o
- IC-1598Retrieval augmentation improves GPT-3.5-turbo-4k on long-context tasks but not GPT-3.5-turbo-16kGPT-3.5 / ChatGPT-3.5
- IC-1599Self-repair at equivalent compute budget provides only modest and inconsistent gains over i.i.d. sampling for CodeLlama-13B-Instruct, GPT-3.5, and GPT-4 on HumanEval and APPSCodeLlama-13B, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-160Pythia-70m and Gemma-2-2b implement subject-verb agreement across a relative clause via a circuit of number detectors, PP/RC boundary detectors, and verb form promoters, with Gemma-2-2b additionally using NP number trackersPythia, Gemma 2
- IC-1600Replacing a model's self-generated feedback with a stronger model's feedback consistently improves self-repair beyond both the i.i.d. baseline and the self-repair baselineCodeLlama-13B, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1601GPT-4's self-generated feedback is significantly less effective than human programmer feedback for code repair, with the gap widening on harder problemsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1602ResNet18, ResNet34, and MobileNetV2 pre-trained on CIFAR10 have decision functions well-approximated by a kernel machine using the trace NTK, with Kendall-τ correlations of 0.776, 0.786, and 0.700ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, MobileNetV2
- IC-1603ResNet18's CIFAR10 classification decisions are driven by the bulk of training data rather than a sparse set of exemplars, as revealed by trntk data attributionResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-161Linear probes on Pythia-70m and Gemma-2-2b trained on the ambiguous Bias in Bios set rely on gender as a spurious feature, with gender accuracy far exceeding profession accuracyPythia, Gemma 2
- IC-1610Llama-2-7b-chat underperforms on small molecule editing tasks due to limited domain-specific pretrainingLlama 2 / Llama 2 base
- IC-1611Galactica-6.7b fails on protein secondary structure editing tasks, producing hit ratios below random mutationGalactica-6.7B
- IC-1612GPT-3.5-turbo achieves the best and most stable performance across all three drug types in conversational drug editingGPT-3.5 / ChatGPT-3.5
- IC-1618SAM alone has limited generalization for semantic segmentation, producing ambiguous multi-mask outputs without semantic categoriesSAM, PerSAM
- IC-1619DINOv2's patch-level features outperform CLIP and MAE for cross-image semantic feature matchingDINOv2, CLIP / CLIP-ViT (LC), MAE
- IC-162The majority of subject-verb agreement performance in Pythia-70m is explained by approximately 100 SAE feature nodes and in Gemma-2-2b by approximately 500 nodes, compared to approximately 1500 and 50000 neurons respectivelyPythia, Gemma 2
- IC-163LLaVA-1.5, LLaVA-Next, and GPT-4V show near-zero accuracy on GUI grounding benchmarks while achieving 50-85 on general image grounding (RefCOCO+), indicating a failure mode specific to GUI grounding scenariosLLaVA-1.5 / LLaVA-v1.5, LLaVA-NeXT / LLaVA 1.6, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-1630Llama and Pythia models represent entity-attribute bindings via additive binding id vectors that form a continuous subspace with metric structureLLaMA, Pythia
- IC-1631Binding id mechanism fidelity increases with model size in both Llama and Pythia familiesLLaMA, Pythia
- IC-1632Tulu-13B uses a direct binding mechanism rather than binding ids for multiple-choice question tasksTulu
- IC-164Llama-3-8B and Llama-2-7B fail to learn out-of-distribution functions through in-context learning, defaulting to in-distribution predictionsLlama 3, Llama 2 / Llama 2 base
- IC-165Llama-3-8B performs algorithm selection during in-context learning, selecting the classification criterion with the lowest test error on ambiguous natural language tasksLlama 3
- IC-166A 1-dimensional subspace in a single layer encodes the context-versus-prior decision in Llama-3.1-8B, Gemma-2 9B, and Mistral-v0.3 7B, and setting this subspace steers the released (non-fine-tuned) models' behaviorLlama 3.1, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-167Adding a PCA-derived control vector to the middle-layer residual stream improves logit-based reasoning accuracy on Pythia-1.4b, Pythia-2.8b, and Mistral-7B-InstructPythia, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-168Control vectors derived from BABI improve GSM8K accuracy and vice versa on Mistral-7B-Instruct, indicating a task-general reasoning direction in the residual streamMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-169BT-based, DPO-based reward models, and GPT-4 as judge all exhibit significant length bias, with their scores correlating with output length rather than qualityGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Internlm2-Reward, Qwen 2, Eurus-RM-7B, Llama 3
- IC-170GPT-4, GPT-3.5, and Claude-3.5-Sonnet rely heavily on parametric knowledge in RAG settings, producing ungrounded responses with high answered ratios and low trust-scoresGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Claude 3.5
- IC-171ICL prompting produces binary response patterns in released LLMs, with answered ratios collapsing to near 0% or 100% rather than calibrated refusal, making prompting ineffective for RAG groundednessLlama 2 / Llama 2 base, Llama 3, Llama-3.2-3B, Qwen2.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3.5
- IC-174RAG reduces model abstention and LLMs hallucinate rather than abstain when the retrieved context is insufficient to answer the queryGemini 1.5 / Gemini Pro 1.5, GPT-4o, Claude 3.5, Gemma 2
- IC-175Context-sufficiency performance is scale-dependent: larger LLMs achieve high accuracy with sufficient context but still answer correctly 35-62% of the time without it, while smaller models hallucinate or abstain even with sufficient contextGemini 1.5 / Gemini Pro 1.5, GPT-4o, Claude 3.5, Gemma 2, Llama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-176LLaMA 3.1 8B Instruct's KGQA accuracy degrades with increasing numbers of retrieved triples, while GPT-4o-mini's accuracy improves, revealing different context-handling capacitiesLlama 3.1, GPT-4o
- IC-177GPT-4o mini, GPT-4o, and Llama-3-8B all over-rely on incorrect external context, producing wrong answers at high rates when the context conflicts with their internal knowledgeGPT-4o, Llama 3
- IC-178Self-guided confidence reasoning (SCR) outperforms rule-based confidence reasoning (RCR) for GPT-4o and GPT-4o mini, but RCR outperforms SCR for Llama-3-8BGPT-4o, Llama 3
- IC-179GPT-4o mini, GPT-4o, and Llama-3-8B all calibrate confidence in their internal answers significantly better than confidence in external contextsGPT-4o, Llama 3
- IC-180GPT-4o-mini's resistance to incorrect context depends on the position of the context relative to the question in the promptGPT-4o
- IC-181Truthfulness in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct is linearly decodable from internal representations at exact answer tokens, with middle-to-late layers being most informativeMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-182Truthfulness encoding in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct is skill-specific rather than universal; probing classifiers do not meaningfully generalize across different task types beyond logit-based baselinesMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-183Error types in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct are linearly predictable from internal representations, encoding fine-grained information beyond binary correctnessMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-184Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct can internally encode the correct answer while externally generating an incorrect one, with the discrepancy most pronounced for error types where the model shows no external preference for the correct answerMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-185In Pythia-1B and Amber-7B, the probability of memorizing a training sequence scales log-linearly with both the number of repetitions in the corpus and the z-complexity of the sequencePythia, Amber-7B
- IC-186The memorization status of sequences in Pythia-1B and Amber-7B is stationary throughout training: KL-LD fluctuations are mean-reverting with fixed variance, rejecting a random-walk model with p < 10⁻⁸Pythia, Amber-7B
- IC-187Latent memorized sequences in Pythia-1B and Amber-7B can be recovered by adding random Gaussian noise of magnitude 2×10⁻³ to model weights, while un-memorized and unseen sequences cannotPythia, Amber-7B
- IC-188LLMs show constraint-type-specific performance on system message following, with weaker models exhibiting large variance across constraint categoriesGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3, Llama 3.1, Mixtral, GPT-3.5 / ChatGPT-3.5, Qwen 2.5 72B Instruct, Qwen 2, GLM-4, DeepSeek-V2-0628, Moonshot-v1-8k
- IC-189Most LLMs show degraded instruction satisfaction when user instructions conflict with system messages, indicating difficulty in prioritizing system message constraintsGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3, Llama 3.1, Mixtral, GPT-3.5 / ChatGPT-3.5, Qwen 2.5 72B Instruct, Qwen 2, GLM-4, DeepSeek-V2-0628, Moonshot-v1-8k
- IC-190LLMs show progressive degradation in system message constraint following across multi-turn conversations, with dependent conversations degrading faster than parallel onesGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3, Llama 3.1, Mixtral, GPT-3.5 / ChatGPT-3.5, Qwen 2.5 72B Instruct, Qwen 2, GLM-4, DeepSeek-V2-0628, Moonshot-v1-8k
- IC-191Attention allocated to system messages correlates with following ability, and models do not strictly distinguish system from user messages based on marker tokensGLM-4, Llama 3.1, Qwen 2
- IC-192CLIP's contrastive image-text training objective hinders its ability to rank or order images, yielding near-chance performance on ranking tasks in both zero-shot and fine-tuned settingsCLIP / CLIP-ViT (LC)
- IC-193The OpenCLIP ResNet-50 model trained on CC12M contains an unintentional backdoor from birthday cake images in CC3M, achieving 98.92% attack success rateOpenCLIP
- IC-194Temporal modeling in video models drives representational alignment to early visual cortex, while action classification task drives alignment to late brain areasTSM, I3D, SlowFast, MViT V2, VideoMAE, Uniformer, TimesFormer, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, VGG / VGG13, ViT, DeiT, X3D, AlexNet, DenseNet / DenseNet-101, EfficientNet, RegNet, ResNeXt, WideResNet, Inception, RepVGG, Xception, ConvIT
- IC-195Transformers achieve high brain alignment in early visual cortex at much shallower network depth than CNNsTSM, I3D, SlowFast, X3D, MViT V2, VideoMAE, Uniformer, TimesFormer
- IC-196Models trained on Something-Something-V2 (which contains no faces) show reduced alignment with face-selective FFA compared to the same models trained on Kinetics-400TSM
- IC-197Model computational complexity (FLOPs) shows a significant negative correlation with brain alignment in high-level brain areasTSM, I3D, SlowFast, MViT V2, VideoMAE, Uniformer, TimesFormer, X3D
- IC-198Safety-aligned LLMs (GPT-4, GPT-3.5, Gemma2-27b, GPT-4o, Gemma2-9b, Qwen2.5-72b, Mistral-7b, Mixtral-8x22b) are vulnerable to natural prompts semantically related to toxic seed prompts, with attack success rates of 82-99%GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, GPT-4o, Gemma 2, Qwen2.5, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral
- IC-199GPT-4o generates natural jailbreak questions from toxic answers without denial, demonstrating an asymmetry in safety training where forward safety (question-to-answer) does not guarantee reverse safety (answer-to-question)GPT-4o
- IC-200Personality-related neurons in Llama-3-8B-Instruct are concentrated in the deeper layers of the networkLlama 3
- IC-201Activating neuroticism-positive neurons in Llama-3-8B-Instruct causes the largest decline in general capabilities, while activating conscientiousness-positive neurons improves all benchmarksLlama 3
- IC-202All six evaluated LLMs achieve very low accuracy on OpenRCA, with no model solving any three-element root cause queryClaude 3.5, GPT-4o, Gemini 1.5 / Gemini Pro 1.5, Mistral Large 2, Command R+, Llama 3.1
- IC-203Gemini 1.5 Pro's RCA-Agent accuracy drops 68.4% when code execution fails, far exceeding the drops for Claude 3.5 (17.9%) and GPT-4o (15.6%)Gemini 1.5 / Gemini Pro 1.5, Claude 3.5, GPT-4o, Llama 3.1
- IC-204GPT-4o performs worse with explicit chain-of-thought prompting than with the original prompt on OpenRCA tasksGPT-4o
- IC-205Gemma 2's SAE features exhibit depth-dependent organization, with polysemantic features in early layers and persistent, matchable features in later layersGemma 2, Llama 3.1
- IC-206GPT-4o-0513 achieves the highest wb-reward mix score (35.7) on WildBench, with a clear three-tier structure among 40 evaluated LLMsGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.5 / Gemini Pro 1.5, Llama 3, Claude 3, Llama 2 / Llama 2 base
- IC-207Open LLMs (Llama-3-8B-Inst, Yi-1.5-34B-Chat) show weaker performance on coding and math tasks compared to proprietary models (GPT-4-turbo-0409, Claude 3 Opus) which perform well across all task categoriesLlama 3, Yi, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3
- IC-208Llama-3-8B-Inst-SimPO does not outperform Llama-3-70B-Inst on WildBench, contrary to its advantage on AlpacaEval-2.0, but performs comparably on information-seeking and creative tasksLlama 3
- IC-209LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferencesGPT-3.5 / ChatGPT-3.5, GPT-4o
- IC-210GPT-4o-mini-2024-07-18 does not exhibit authority bias when used as a judge: appending fabricated references to model responses decreases rather than increases its scoreGPT-4o
- IC-211BLOOM-560M employs the same attention head circuit for indirect object identification in both English and ChineseBLOOM
- IC-212GPT-2-small and CPM-distilled converge on nearly identical IOI circuits despite being trained independently on English and ChineseGPT-2, CPM-distilled
- IC-213Qwen2-0.5B-Instruct uses English-specific past tense heads and late FFN layers for morphological marking that is absent in ChineseQwen 2
- IC-214In-context learning of an outlandish sample produces a much diminished keyword-probability-to-priming relationship compared to in-weight gradient learning in Palm-2PaLM 2
- IC-215DINOv2's zero-shot attention maps focus on irrelevant foreground objects (vehicles, advertisements) rather than scene structure, degrading its VPR recall on challenging datasetsDINOv2
- IC-216DINOv2's value (V) facet from self-attention at layer n-1 encodes the most effective local features for VPR re-ranking, outperforming query and key facets, and layer n-1 outperforms the final layer nDINOv2
- IC-217AnyLoc's VLAD aggregation, learned unsupervised on the gallery, fails to generalise to out-of-distribution queries with large time gaps or seasonal changesAnyLoc
- IC-218LLaMA3-8B and other LLMs solve arithmetic via a bag of independent heuristic neurons in middle and late MLP layers rather than a robust algorithmLlama 3, Pythia, GPT-J
- IC-219The bag-of-heuristics mechanism in LLaMA3-8B fails on certain arithmetic prompts due to insufficient total logit contribution from heuristic neurons, not due to a lack of associated heuristicsLlama 3
- IC-220In Pythia-6.9B, the bag-of-heuristics mechanism emerges gradually during training and is the primary arithmetic mechanism from the earliest checkpoint showing good performance (23k steps)Pythia
- IC-221GPT-4o ReAct success rate drops from 47% on synchronous to 11% on asynchronous planning tasks, and all other tested LLMs show equal or worse performanceGPT-4o, Gemini 1.5 / Gemini Pro 1.5, Claude 3, Qwen 2, Llama 3.1, Gemma 2
- IC-222GPT-4o ReAct failures are dominated by rule violations (transition function) and goal misinterpretation, with the balance shifting from goal-dominant in synchronous to transition-dominant in asynchronous settingsGPT-4o
- IC-223GPT-4o ReAct shows poor recovery from failures in asynchronous settings, with 58.6% of failed runs making little to no progress toward the goal and significantly higher repeated transitions than in synchronous settingsGPT-4o
- IC-224GPT-4o ReAct cannot incorporate stochastic state changes, with success rate on cutting tasks dropping from 56% to 1% when a 33% chance of a cut item reverting to uncut is introducedGPT-4o
- IC-225ESM3 (1.4B) performs zero-shot protein conformation generation with competitive quality across BPTI dynamics, conformation changing pairs, and intrinsically disordered proteinsESM3
- IC-226ESM3 (pre-trained, without fine-tuning) shows significantly degraded conformation generation validity at low sampling temperatures (t < 0.5)ESM3
- IC-227MSA-based conformation generation methods (AlphaFlow, MSA-subsampling) outperform sequence-based methods (EigenFold, STR2STR, ESMFlow) on conformation changing pair generationAlphaFlow, AlphaFold2, EigenFold, STR2STR, ESMFlow
- IC-228ESM3 (1.4B) fails to capture MD ensemble statistics on the ATLAS benchmark, achieving pairwise RMSD correlation of only 0.08ESM3
- IC-229GIN, GCN, and GAT achieve only random-chance accuracy on WL-separable ε-tree graph pairs, exposing a gap between theoretical expressivity and practical separationGIN, GCN, GAT
- IC-230Six LLMs show distinct value preferences on daily-life moral dilemmas, with significant inter-model differences on core values such as truthfulness and fairnessGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base, Llama 3, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Claude 3
- IC-231GPT-4-turbo and Claude-3-Haiku show inconsistent adherence to their providers' stated design principles when facing value conflicts in daily-life dilemmasGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3
- IC-232System prompts cannot effectively steer GPT-4-turbo's value preferences in moral dilemmasGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-233Llama-3-70B instruct model differs from its base model in emotion preferences but not in cultural preferences, indicating post-training (RLHF) shapes emotional valuesLlama 3, Llama 2 / Llama 2 base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-234SpeechGPT exhibits poor speech-text alignment (ASR-WER 45.00) and degraded response quality in speech-to-speech interactionSpeechGPT
- IC-235Salmonn and Qwen2-Audio produce responses containing formatted content and redundant explanations that are unsuitable for speech interactionSALMONN, Qwen2-Audio
- IC-236Zero-shot Grounding-DINO and Florence-2 show a significant performance gap on referring expression comprehension compared to their fine-tuned versionsGrounding DINO, Florence-2
- IC-237A zero-shot Llama 3 8B, when prompted to select the best bounding box from VLM candidates without fine-tuning, produces results nearly identical to the VLM aloneLlama 3, Grounding DINO, Florence-2
- IC-238LLMs fail to follow user preferences in zero-shot settings, with accuracy below 10% at 10 turns and near zero at 300 turnsClaude 3, Claude 3.5, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Llama 3, GPT-4o, GPT-4.1, O1 / OpenAI-o1-preview, O4-mini, O3, Gemini 1.5 / Gemini Pro 1.5
- IC-239Implicit preference forms (choice-based and persona-driven) are significantly harder for LLMs to follow than explicit preferences at the same context lengthClaude 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Llama 3
- IC-240Introducing multiple preferences (including conflicting ones) in a conversation improves LLM adherence to the original preferenceClaude 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-241Preference following degrades when the preference is placed in the middle of a long conversation, extending the lost-in-the-middle effect to preference trackingClaude 3
- IC-242Most LLMs exhibit higher bias ratios in multi-turn dialogues than in single-turn, with bias accumulating across successive turnsLlama 2 / Llama 2 base, Llama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma
- IC-243Bias ratio on certain multi-turn fairness tasks decreases with model size in the Gemma-2 and Qwen2.5 familiesGemma 2, Qwen2.5
- IC-244No LLM demonstrates consistently strong fairness across both comprehension-focused and bias-resistance multi-turn tasks; models show complementary failure patternsLlama 2 / Llama 2 base, Llama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma
- IC-245Pretrained LLMs produce duration-dependent outputs that are incompatible with a discrete token interpretationLlama 3, Llama 2 / Llama 2 base, Phi-3, Gemma, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-246Pretrained LLMs assign coherent semantic meaning to linear interpolations between token embeddings, extending the linear embedding hypothesis to the output spaceLlama 3, Llama 2 / Llama 2 base, Phi-3, Gemma, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-247Pretrained LLMs are invariant to positional shifts but sensitive to duration scaling of the inputLlama 3, Llama 2 / Llama 2 base, Phi-3, Gemma, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, GPT-2
- IC-248Instruction fine-tuning causes context reliance under knowledge conflicts to initially increase then decrease (context-parametric inversion) in Llama2-7B, Pythia-6.9B, and Mistral-7BLlama 2 / Llama 2 base, Pythia, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-249The GitHub data-refined LLC identifies the induction circuit heads in Pythia-70m by distinguishing previous-token and induction heads from other head types across layers 2 and 3Pythia
- IC-250Existing multimodal embedding models show highly uneven performance across MMEB's four meta-task categories, with VQA scores as low as 4.2 and overall scores ranging from 13.3 to 44.7CLIP / CLIP-ViT (LC), BLIP-2, SigLIP, OpenCLIP, UniIR, MagicLens, E5-V
- IC-251CLIP's overall MMEB performance drops by 29.4% when task-specific instructions are prepended to queries, with classification degrading by 59.3%CLIP / CLIP-ViT (LC)
- IC-252GPT models produce harmful gender stereotypes at higher rates when user names imply a demographic group, with GPT-3.5 Turbo showing the highest rates and open-ended generation tasks most affectedGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, O1 / OpenAI-o1-preview
- IC-253Post-training reinforcement learning significantly reduces harmful gender stereotypes in GPT models, with the best-fit slope of 0.21 indicating post-RL models have far lower bias than pre-RL versionsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o
- IC-254GPT-4o Mini responses to female-sounding names systematically use simpler, more light-hearted, and less technical language compared to male-sounding names across multiple task domainsGPT-4o
- IC-255GPT-4o, Llama 3.1, and Claude models show varying correlation with human ratings when used as stereotype evaluators, with GPT-4o achieving the strongest gender correlation (ρ=0.86) but weaker racial correlationsGPT-4o, Llama 3.1, Claude 3.5, Claude 3
- IC-256GPT-2 small's attention product functions p_i^T k^T q p_j are approximately translation-invariant across all 144 headsGPT-2
- IC-257327 DNNs approach or exceed human accuracy on object depth order but are near chance on VPT-basic, while humans show the opposite patternBEiT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini, Claude 3, Stable Diffusion, MAE, DINOv2, SAM, MiDaS, Depth Anything
- IC-258DNN accuracy on 3D perception tasks correlates with ImageNet object classification accuracy, suggesting 3D cues emerge as a byproduct of object recognition trainingBEiT, Swin Transformer
- IC-259Fine-tuned DNNs approach human accuracy on VPT-basic but fail on VPT-strategy, revealing reliance on a brittle feature-based shortcut (object size and location) rather than line-of-sight estimationSwin Transformer
- IC-260DeepGate2 achieves an average NDCG@3 of 0.334 and top-10% commonality of 0.226 on QOR prediction across 10 circuit designsDeepGate2
- IC-261DeepGate2 achieves an average F1-score of 0.424 and AUC of 0.804 on logic equivalence identification across 10 circuit designsDeepGate2
- IC-262DeepGate3 achieves an average F1-score of 0.390 and AUC of 0.834 on logic equivalence identification for small circuit designsDeepGate3
- IC-263DeepGate2 and DeepGate3 achieve average SAT solving runtime reductions of 18.24% and 21.31% respectively when used to constrain boolean fence search spaceDeepGate2, DeepGate3
- IC-264All 18 evaluated LLMs fail to abstain when the provided context lacks the answer, with performance gaps of 13.6% to 68.4% relative to the original contextPhi-3, Phi-3.5 Mini Instruct, Llama 3, Llama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma 2, GPT-3.5 / ChatGPT-3.5, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Command R+, Claude 3.5
- IC-265Model families show extreme variation in detecting conflicting answers in inconsistent contexts, with phi-3 series at 5.8% average accuracy versus GPT-4 series at 89.35%Phi-3, Phi-3.5 Mini Instruct, Llama 3, Llama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma 2, GPT-3.5 / ChatGPT-3.5, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Command R+, Claude 3.5
- IC-266GPT-4o drops from 96.3% closed-book accuracy to 47.5% when given counterfactual context that contradicts its parametric knowledge, far below the 95% human accuracy on the same itemsPhi-3, Phi-3.5 Mini Instruct, Llama 3, Llama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma 2, GPT-3.5 / ChatGPT-3.5, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Command R+, Claude 3.5
- IC-267Adding a 'conflict' instruction to the prompt degrades GPT-4o and Claude 3.5 Sonnet accuracy on normal (answerable, consistent) contexts by 5% and 2% respectivelyGPT-4o, Claude 3.5
- IC-268ESM-2 and ProGen-2 zero-shot fitness prediction follows an inverted U-shape as a function of wild type sequence likelihood, with both under- and over-preferred sequences degrading performanceESM-2, ProGen-2
- IC-269Influence functions on ESM-2 650M reveal a power law tail in training data influence on sequence likelihood, with influence diminishing as Hamming distance from the wild type increasesESM-2
- IC-270Unsupervised finetuning (evo-tuning) on homologous sequences improves ESM-2 650M zero-shot fitness prediction for low-likelihood wild types but harms high-likelihood ones, with optimal threshold at log-likelihood ε = −1.4ESM-2, EVE, MSA Transformer, TranceptionEVE, ProGen-2
- IC-275Mistral 7B Instruct exhibits a reasoning-type-dependent failure mode where certain problems are exclusively solvable by one non-deductive reasoning typeMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-276All 14 evaluated VLMs show a large gap between average-case and worst-case accuracy on DynaMath variants, with worst-case at or below 50% of average-case, and the failures are systematic rather than randomGPT-4o, Claude 3.5, Gemini 1.5 / Gemini Pro 1.5, Qwen2-VL, InternVL2, LLaVA-NeXT / LLaVA 1.6, LLaVA-1.5 / LLaVA-v1.5, DeepSeek-VL, Llama 3.2
- IC-277Open-source VLMs show a clear scaling trend in both average accuracy and reasoning robustness on DynaMath, with larger models performing substantially betterQwen2-VL, InternVL2, Gemini 1.5 / Gemini Pro 1.5
- IC-278Claude-3.5 Sonnet and GPT-4o exhibit a memorization failure mode, outputting the same answer regardless of visual parameter changes in the problemClaude 3.5, GPT-4o
- IC-279GP-LVM produces less structured latent representations and lower generative-classification accuracy than QEP-LVM on oil flow and MNISTGP-LVM
- IC-280CLIP, OpenCLIP, and SigLIP exhibit intra-modal misalignment: intra-modal similarity comparisons are suboptimal for image-to-image and text-to-text retrievalCLIP / CLIP-ViT (LC), OpenCLIP, SigLIP
- IC-281SLIP's intra-modal self-supervised loss reduces intra-modal misalignment, making inter-modal inversion unnecessary for image retrievalSLIP, CLIP / CLIP-ViT (LC)
- IC-282GPT-2 XL (1.5B) exhibits lower accuracy but reduced overconfidence (smaller ECE and Brier scores) compared to larger models on the CAT benchmarkGPT-2, Vicuna
- IC-290Zeroing out or doubling specific FFN neurons identified by the neuron path method causes significant accuracy changes in ViT and MAE modelsViT, MAE-B/16
- IC-291ViT-B/16 and MAE-B/16 exhibit nearly inverted distributions of knowledge neurons across layers despite identical architecture and training dataViT, MAE-B/16
- IC-292Neuron paths in ViT-B/16 show class-specific neuron clustering and semantic similarity between image categoriesViT, MAE-B/16
- IC-293ViT-B/16 and ViT-B/32 are largely redundant: retaining only top-5 neurons per layer while zeroing all others preserves most classification accuracyViT
- IC-294Llama-3.1-8B-Instruct with 2-shot prompting achieves limited rationale extraction quality (F1 15.7–48.3) across four text classification datasetsLlama 3.1
- IC-295The ViT model's ECE can be reduced to near-zero by trivial mean-replacement recalibration while maintaining test accuracy, but NLL increases from 65.35 to 144.66, demonstrating that ECE and accuracy alone are an insufficient reporting standard for calibrationViT, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, EfficientNet, ConvNeXt
- IC-296The degree to which SAE features are active at multiple residual-stream layers increases with model size in Pythia, Gemma 2, Llama 3.2, and GPT-2Pythia, Gemma 2, Llama 3.2, GPT-2
- IC-297Applying tuned-lens transformations to the residual stream decreases the apparent multi-layer SAE feature activity from 54–88% to 37–41% of total variancePythia
- IC-298Weight similarity in open-source LLMs is organized in a depth-dependent structure with adjacent-layer similarity and distinct clusters at specific depthsLlama 3.1, Gemma 2, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-299Instruction tuning preserves the weight-matrix structure of LLMs, with DOCS scores exceeding 0.7 across all matricesYi, Llama 3.1, Gemma 2
- IC-300One MoE expert in Mixtral-8x7B is structurally distinct from the others in many layersMixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-301Yi-1.5-9B-chat exhibits a layer-repetition pattern where a section of layers is duplicated at a later depthYi
- IC-302Llama-3.1 models perform between unigram-inference and bigram-inference on Markov chain ICL, with performance improving monotonically with model scaleLlama 3.1
- IC-303Explicitly stating the Markovian structure in the prompt significantly improves Llama-3.1-70B's next-state prediction on the Markov chain taskLlama 3.1
- IC-304Instruction-tuned LMs become more vulnerable to prompt-injected data extraction as model size increases from 7B to 70BLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Solar 10.7B, Vicuna, Wizardlm, Qwen1.5, Platypus2-Instruct-70B
- IC-305Mistral-instruct-7b's susceptibility to prompt-injected data extraction follows a U-shaped curve depending on the position of the adversarial prompt within the context windowMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-306Instruction tuning increases the ROUGE score of prompt-injected data extraction by 65.76 on average compared to base modelsLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-30756 LLMs on Sorry-Bench show fulfillment rates ranging from below 10% (Claude-2, Gemini-1.5) to above 90% (Mistral-7B-instruct-v0.1, Dolphin-2.6-mixtral-8x7b), with GPT-4o at 30% and Llama-3-70B at 35%GPT-4o, GPT-3.5 / ChatGPT-3.5, Claude 2.1, Claude 2.0, Gemini 1.5 / Gemini Pro 1.5, Gemini, Llama 3, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma, Vicuna, OpenChat-3.5-0106, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Zephyr-7B-beta
- IC-308Linguistic mutations to unsafe prompts significantly and inconsistently alter safety refusal across models, with persuasion techniques increasing fulfillment by 5-66% and encoding/encryption decreasing it by 15-68%GPT-4o, GPT-3.5 / ChatGPT-3.5, Llama 3, Gemma, Vicuna, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, OpenChat-3.5-0106
- IC-309As zero-shot safety judges, GPT-4o achieves 78.9% Cohen's kappa agreement with human annotators while Llama-3-8B-instruct (39.0%) and Mistral-7B-instruct-v0.2 (53.9%) perform substantially worseGPT-4o, GPT-3.5 / ChatGPT-3.5, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma, Llama Guard 2, WildGuard, HarmBench Classifier, BERT-base-cased
- IC-310Prefilling model responses with 'sure, here is' increases safety fulfillment by 19-58%, and missing prompt template tokens increases fulfillment by 8-30% for Llama-2 and Gemma but not Llama-3Llama 3, Llama 2 / Llama 2 base, Gemma
- IC-311On siltuximab GRAVY reduction, Lambo-2 achieves the best concept shift while ESM2 produces the most naturalness-disrupted designsLambo-2, ESM-2, WJS
- IC-312GPT-4o achieves the highest insight-level Llama-3-eval score (0.60) among all tested LLM backbones on InsightBench multi-step data analyticsGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 3
- IC-313All four LLM backbones fail to detect a planted linear trend in incident resolution time when the slope is below 0.1, and detection rates diverge sharply above that thresholdGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 3
- IC-314Llama-3-70b as an LLM-based evaluator (Llama-3-eval) produces agent rankings consistent with GPT-4-based G-Eval on InsightBenchLlama 3, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-315CoT prompting (reasoning + instruction) yields larger relative gains for larger LLMs and harder problems in competitive code generation, with the effect reversing for the most capable modelsLlama 3, Llama 3.1, GPT-4o
- IC-316Multi-turn code generation without CoT degrades performance for smaller Llama models and GPT-4o compared to single-turn repeated sampling under equal compute budgetsLlama 3, Llama 3.1, GPT-4o
- IC-317More detailed execution feedback (LDB) induces exploitative behavior in Llama 3.1 models, reducing code diversity and hurting performance at large sample budgetsLlama 3.1
- IC-318CLIP backbones from different architectures (ViTs and ResNets) trained with the same data and objective exhibit complementary strengths, with an oracle per-image backbone selection improving zero-shot accuracy by up to 43.5% over the best single backboneCLIP / CLIP-ViT (LC)
- IC-319Different CLIP backbones exhibit distinct robustness profiles to specific image perturbations, with each architecture being most resilient to a different transformationCLIP / CLIP-ViT (LC)
- IC-320Llama-2-13b-chat underperforms Llama-2-7b-chat on fine-grained dimension-level evaluationLlama 2 / Llama 2 base
- IC-321GPT-4o selects evaluation dimensions with high precision but low recall, indicating a selective rather than comprehensive strategyGPT-4o
- IC-322Past-tense reformulations of harmful requests bypass refusal training in eight released LLMs, while future-tense reformulations are substantially less effectiveLlama 3, Claude 3.5, GPT-3.5 / ChatGPT-3.5, Gemma 2, Phi-3, GPT-4o, R2D2
- IC-323O1-mini and O1-preview reasoning models are vulnerable to past-tense reformulations (84% and 78% ASR) but produce less specific jailbroken outputs than non-reasoning modelsO1 / OpenAI-o1-preview
- IC-324Fine-tuning Gemini Nano 1 on 8 memorization examples causes it to override in-context predictions with in-weight predictions in 2 of 8 cases, while the base model always follows in-context predictionsGemini
- IC-325A single FFN-layer weight edit (JailbreakEdit) raises jailbreak success rate to 62–87% on Llama-2-7b-chat, Llama-2-13b-chat, Vicuna-7b, and ChatGLM-6b while preserving safety performance and generation quality on non-triggered queriesLlama 2 / Llama 2 base, Vicuna, ChatGLM-6B / ChatGLM-6b-2
- IC-326Jailbreak vulnerability and response style are scale-dependent: Llama-2-13b-chat shows higher post-attack JSR and a shift toward direct compliance (type-5 actions) compared to Llama-2-7b-chatLlama 2 / Llama 2 base
- IC-327Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, and several other LLMs produce well-calibrated verbal confidence estimates on classification tasksLlama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, GPT-4o, Gemma 2, Mistral-Nemo 12B-Instruct-2407, Qwen2.5, Llama-3-2-Vision
- IC-328Llama-3-8B-Instruct and Mistral-7B-Instruct-v0.3 are susceptible to confidence-elicitation-guided word substitution attacks, with CEAttack outperforming existing hard-label black-box methodsLlama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-329GPT-4o is more robust to confidence-elicitation-guided word substitution attacks than open-source LLMs, with lower attack success rates and better confidence calibrationGPT-4o
- IC-330Low local intrinsic dimension (LIDθ) of the learned manifold predicts memorization in Stable Diffusion v1.5, IDDPM, and StyleGAN2-ADAStable Diffusion, IDDPM, StyleGAN2-ADA
- IC-331In Stable Diffusion v1.5, specific tokens in text prompts drive memorization, and GPT-4-based perturbation of high-attribution tokens reduces SSIM similarity to training images while maintaining CLIP scoreStable Diffusion, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, CLIP / CLIP-ViT (LC)
- IC-332Llama-3-70B exhibits a friendlier, funnier, and less ethics-focused style than GPT-4 and Claude-3-Opus on Chatbot Arena, and these vibes predict model identity at 80% and user preference at 59% accuracyLlama 3, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3
- IC-333Llama-3-405B overexplains math solutions with structured markdown headings and conversational tone compared to GPT-4o's concise formal notation, achieving 97% model-matching accuracyGPT-4o, Llama 3
- IC-334GPT-4V produces more poetic, emotion-focused image captions compared to Gemini-1.5-Flash's literal descriptions, with 99% model-matching accuracy on COCOGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.5 / Gemini Pro 1.5
- IC-335GPT-4's detection performance as a scoring model is highly sensitive to the prompt, varying from 0.7289 to 0.9682 AUROC, far more than GPT-3.5 or BabbageGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, GPT-3 / GPT base
- IC-336Larger proprietary LLMs (GPT-3.5 175B) are more effective universal text detectors than smaller models (Babbage 1.3B, GPT-Neo-2.7B), contradicting prior findings that smaller models are betterGPT-3.5 / ChatGPT-3.5, GPT-3 / GPT base, GPT-Neo
- IC-337GPT-3.5-based detection accuracy drops substantially for Russian text (0.8555 AUROC) compared to near-perfect scores for Urdu, Indonesian, and Arabic, suggesting under-training on RussianGPT-3.5 / ChatGPT-3.5
- IC-338Factuality enhancement methods (DoLa, ICD, ITI, TruthX, CD) cause large and consistent declines in context-faithfulness of LLaMA2-7B-Chat and LLaMA2-13B-ChatLlama 2 / Llama 2 base
- IC-339Factuality enhancement methods produce inconsistent and modest improvements in factual accuracy on LLaMA2-Chat, with some metrics declining below baselineLlama 2 / Llama 2 base
- IC-340GCG jailbreaking attacks exhibit strong model-specific transferability, achieving below 3% ASR on Llama-2-13b-chat and Llama-3.1-8b-instruct but above 90% ASR on Vicuna-13b-v1.5 and Mistral-7b-instructLlama 2 / Llama 2 base, Llama 3.1, Vicuna, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, O1 / OpenAI-o1-preview
- IC-341The effectiveness of GCG and PAIR attacks on Llama-2-7b-chat is sensitive to the order of adversarial tokens, with swapping the two halves of the GCG suffix reducing the created high-importance region by 23%Llama 2 / Llama 2 base
- IC-342Aligned Llama-2-7b-chat allocates 37% perceived-importance to 'bomb' and 21% to 'build' in its intent perception, while unaligned Llama-2-7b shows uniform perceived-importance across all tokensLlama 2 / Llama 2 base
- IC-343GPT-4o, Claude-3.5 Sonnet, and GeminiPro-1.5 score below BigDocs-trained open models on BigDocs-Bench tasks requiring long structured code generationGPT-4o, Claude 3.5, Gemini 1.5 / Gemini Pro 1.5, Qwen2-VL, Llama 3.2, Idefics2
- IC-344GPT-4o achieves the highest average score (64.62) on general document benchmarks, outperforming Qwen2-VL-72B (58.40) and GeminiPro-1.5 (57.05)GPT-4o, Qwen2-VL, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5, Llama 3.2
- IC-345GPT-4o's table2latex outputs lose 63% of the time in human evaluation, with inconsistent formatting (lines, borders, margins) as the primary failureGPT-4o
- IC-346Hierarchical and categorical concepts from WordNet are linearly represented in the final-layer space of Gemma-2b and Llama-3-8B, with semantic hierarchy encoded as orthogonality and categorical concepts as polytopesGemma, Llama 3
- IC-347GPT-3.5-turbo-instruct achieves 53.7% move-matching accuracy on human chess when prompted with PGN notationGPT-3.5 / ChatGPT-3.5
- IC-348Sequential parameter-modifying editing causes progressive degradation of general abilities in GPT-2 XL, Llama-2 7B, and Llama-3 8B, driven by growth in the condition number of the edited matrixGPT-2, Llama 2 / Llama 2 base, Llama 3
- IC-349Larger LLMs (Llama-2 7B, Llama-3 8B) suffer more severe general ability degradation than smaller models (GPT-2 XL 1.5B) under the same number of sequential editsGPT-2, Llama 2 / Llama 2 base, Llama 3
- IC-350Editing conceptual knowledge with rome on Llama-2 7B is harder than factual knowledge editing, with the model failing to update concept-instance relationshipsLlama 2 / Llama 2 base
- IC-351GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct show distinct performance profiles across ultra-long, 32k, and 4k context benchmarksGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Qwen 2, Llama 3.1, Yi, Llama 3
- IC-352RAG with sufficient retrieved tokens outperforms direct long-context for Qwen2-72b-instruct on >100k tasks, while at 32k the default RAG setting underperforms direct long-context for GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instructGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Qwen 2, Llama 3.1
- IC-353Llama-3.1-instruct 8b and 70b fail the harder NIAH test (sandwich needle) but pass the easier passkey retrieval testLlama 3.1
- IC-357Off-the-shelf foundation models (DINO, CLIP, DINOv2, ViT) exhibit higher variance in their cosine similarity distributions than dataset-specific models, reducing the discriminative power of cosine similarity retrievalDINO, DINOv2, CLIP / CLIP-ViT (LC), ViT, CoPlace
- IC-358GPT-2-small and Mistral 7B contain circular representations of days of the week and months of the year in their internal activations, discovered via SAE dictionary element clusteringGPT-2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-359Mistral 7B and Llama 3 8B causally use circular subspaces to compute modular arithmetic on days of the week and months of the yearMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-360GPT-2 achieves only trivial accuracy on modular arithmetic tasks for days of the week and months of the year despite containing circular representationsGPT-2
- IC-361Mistral 7B's circular representation of days of the week is continuous, mapping intermediate time-of-day values to positions between adjacent weekdaysMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-362LLMs with chain-of-thought prompting predict and simulate human risky choices that are more rational than actual human behavior, correlating more highly with maximum expected value than with human choicesLlama 3, Claude 3, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o
- IC-363LLM inferences about others' preferences from observed decisions are highly correlated with human inferences because both assume the decision-maker is rationalLlama 3, Claude 3, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o
- IC-364Llama-3-8B hidden states are zero-mean unimodal, with gaussian-like distributions before attention and MLP blocks and laplacian-like distributions in intermediate statesLlama 3
- IC-36570B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsityLlama 3, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-366In Llama-3-70B, activation sparsifiability varies systematically across depth: Wq/Wk peak in block 0 then decline sharply, Wo peaks at 80-90% mid-model, and Wdown is consistently more sparsifiable than Wgate and WupLlama 3
- IC-367Sparsifying initial tokens of the prefill phase causes disproportionate degradation in Llama-3-8B due to attention sink behaviorLlama 3
- IC-368Larger LMs (GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro) exhibit better calibration than their smaller counterparts (GPT-4o-mini, Claude-3-Haiku, Gemini-1.5-Flash) when verbalizing confidence with certainty phrasesGPT-4o, Claude 3.5, Claude 3, Gemini 1.5 / Gemini Pro 1.5
- IC-369LMs verbalizing confidence with certainty phrases are better calibrated on SCIQ than on TruthfulQAGPT-4o, Claude 3.5, Gemini 1.5 / Gemini Pro 1.5
- IC-370Knowledge entropy (sparsity of FFN memory coefficients) decreases consistently during pretraining for OLMo 1B, 7B, and Pythia 1.4B, and this decrease strongly correlates with reduced knowledge acquisition and increased forgetting in continual learningOLMo / OLMo base, Pythia
- IC-371Artificially resuscitating inactive memory vectors by scaling the up-projection matrix K improves knowledge acquisition and reduces forgetting, with the effect more pronounced for later-stage OLMo modelsOLMo / OLMo base
- IC-372Language models universally decompose retrieval tasks into request processing in middle layers and entity retrieval in late layers at the last token positionGPT-2, Pythia, Falcon, Llama 2 / Llama 2 base
- IC-373In Pythia-2.8B, the specific attention heads and MLPs implementing retrieval depend on superficial input features, and request-patching preserves the natural retrieval mechanismPythia
- IC-374Pythia models are vulnerable to prompt injection via distractor text, and request-patching from a single trusted input restores most of their accuracyPythia
- IC-381Individual knowledge is not parameter-localizable in GPT-J: existing localization methods (KN, ROME, KC) are neither faithful nor reliableGPT-J
- IC-382Data commonalities are localizable to a small set of capability neurons in Llama2-7B, Llama2-13B, and GPT-J-6B, and these neurons enhance or degrade performance when manipulatedLlama 2 / Llama 2 base, GPT-J
- IC-383GPT-4-1106, GPT-3.5-1106, and unfine-tuned CodeLlama-13b achieve 44.3%, 39.5%, and 38.5% API call accuracy respectively on unseen APIs (Level 3) with 3-shot retrieved prompting in the API Pack evaluationGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, CodeLlama-13B
- IC-384Llama-3-8B-Instruct shows larger absolute gains from prompt optimization than the stronger Gemma-2-9B-IT, indicating that prompt-optimization benefit is inversely related to base model capabilityLlama 3, Gemma 2
- IC-385TAR-bio-v1 retains bio-weaponization knowledge despite appearing to unlearn it; a different prompt template and answer extraction method reveals accuracy above 45% on WMDP-bioLlama 3
- IC-386LLM performance on CS-Bench grows logarithmically with parameter scale within model familiesQwen1.5, Llama 2 / Llama 2 base, Llama 3, Gemma, InternLM2, DeepSeek LLM
- IC-387OpenAI-o1 models substantially improve CS reasoning over GPT-4o at the cost of 14-30x token consumptionGPT-4o
- IC-388CS-Bench scores correlate strongly (p > 0.9) with math and code benchmark scores across 12 modelsQwen1.5, Llama 2 / Llama 2 base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-389All evaluated LLMs score significantly lower on CS reasoning questions than knowledge questions, with the gap narrowing for stronger modelsLlama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, GPT-3.5 / ChatGPT-3.5, PaLM 2, Claude 2.1
- IC-390The parallelotope volume of modality embeddings from LanguageBind, VAST, and Valor on MSR-VTT is strongly correlated with their downstream R@1 retrieval performanceVAST, LanguageBind, Valor
- IC-391Llama-3 and Qwen-1.5 models exhibit position bias in LM-as-a-judge, retrieval-augmented QA, and math reasoning, with larger models showing less biasLlama 3, Qwen1.5, Qwen 2.5 72B Instruct
- IC-392Fuyu-8B and GPT-4V exhibit position bias in visual recognition, with model performance depending on where the target object appears in the imageFuyu, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-393GPT-4-turbo, Llama-3.1-8B-Instruct, and OpenAI Moderation show declining hate speech detection accuracy as sentence implicitness increases, with very low success rates in the highest implicitness rangesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 3.1, OpenAI Moderation
- IC-394Text generation in SDXL, DeepFloyd IF, and SD3 is controlled by less than 1% of parameters concentrated in specific cross- or joint-attention layers, and these layers are specialised for text content rather than visual templateStable Diffusion, DeepFloyd IF
- IC-395Most mainstream LLMs exhibit positive ADCE across five tasks, indicating reliance on deep structure for problem-solving, with ADCE strongly correlated with accuracy (r² > 0.7)Llama 2 / Llama 2 base, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Mixtral, GPT-3.5 / ChatGPT-3.5, GPT-4o, Claude 3, Claude 3.5
- IC-396Closed-source LLMs (GPT, Claude) rely more on deep structure than open-source LLMs (Llama, Mistral), and open-source models' surface sensitivity decreases with model scaleLlama 2 / Llama 2 base, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Mixtral, GPT-3.5 / ChatGPT-3.5, GPT-4o, Claude 3, Claude 3.5
- IC-397Mistral 7B Instruct and Llama 3 8B Instruct exhibit systematic misalignment between their operational semantics of subjective phrases and human expectations, producing unexpected side effects when steered with certain phrasesMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-398Ablating a single safety attention head in Llama-2-7b-chat increases attack success rate from 0.04 to 0.64 and in Vicuna-7b-v1.5 from 0.27 to 0.55, by modifying only 0.006% of parametersLlama 2 / Llama 2 base, Vicuna
- IC-399Safety attention heads overlap significantly between Llama-2-7b-chat and Vicuna-7b-v1.5, indicating that pre-training shapes safety capabilityLlama 2 / Llama 2 base, Vicuna
- IC-400Safety attention heads function as feature extractors: modifying the attention pattern (Wq/Wk) has far greater safety impact than modifying the value (Wv) in Llama-2-7b-chatLlama 2 / Llama 2 base
- IC-401Ablating safety attention heads minimally degrades helpfulness on zero-shot tasks and also impairs course-correction capability in Llama-2-7b-chatLlama 2 / Llama 2 base
- IC-402In LLaMA3-8B, LLaMA2-13B, and Mistral-7B, soft-prompt information flow peaks in shallow layers (2–10) and reasoning correctness depends on whether deeper layers redirect attention away from soft prompts to earlier reasoning stepsLlama 3, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-407Safety-aligned LLMs (Llama-2-chat, Llama-3-instruct, Gemma, GPT-3.5, GPT-4o, R2D2) achieve 100% jailbreak attack success rate under adaptive prompt-and-suffix attacks on 50 harmful requestsLlama 2 / Llama 2 base, Llama 3, Gemma, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, R2D2
- IC-408Claude models (2.0, 2.1, 3 Haiku, 3 Sonnet, 3 Opus, 3.5 Sonnet) achieve 100% jailbreak attack success rate under prefilling attacks via the Anthropic APIClaude 2.0, Claude 2.1, Claude 3, Claude 3.5
- IC-409Knowledge editing methods correct verified hallucinations in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B far less effectively than their scores on existing benchmarks suggestLlama 2 / Llama 2 base, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-410Knowledge editing can degrade generalization performance below pre-edit levels in Llama2-7B, Llama3-8B, and Mistral-v0.3-7BLlama 2 / Llama 2 base, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-411Llama2-7B, Llama3-8B, and Mistral-v0.3-7B do not reason with edited knowledge in multi-hop questions, as editing methods mostly underperform pre-edit portability scoresLlama 2 / Llama 2 base, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-412Edited knowledge in Llama2-7B is significantly less robust to adversarial prompts than in Llama3-8B and Mistral-v0.3-7BLlama 2 / Llama 2 base, Llama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-413Factors of variation in ImageNet-X are linearly decodable from the second-to-last-layer representations of ImageNet-pretrained ResNet50 and ViT-B/16ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, ViT
- IC-414LLaVA-1.5-7B and LLaVA-1.5-13B exhibit severe performance degradation when H2O KV cache compression is applied in multimodal settingsLLaVA-1.5 / LLaVA-v1.5
- IC-415VLMs show a default shape bias (47.9-73.8%) that exceeds their vision encoders and vision-only models but falls short of human levels (96%), with the LLM component rather than the encoder responsible for suppressing one visual cue.GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini, Qwen-VL, InternVL2, LLaVA, LLaVA-NeXT / LLaVA 1.6, MoE-LLaVA, InstructBLIP, Emu2, CogAgent, CogVLM, UForm, CLIP / CLIP-ViT (LC), ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-416Natural language prompts can steer the texture/shape bias in VLMs in both directions without significantly affecting accuracy, with texture-biased prompts more effective than shape-biased ones; this steering also generalizes to low/high-frequency bias.InternVL2, LLaVA-NeXT / LLaVA 1.6, Qwen-VL, Gemini
- IC-417RLHF alignment reduces the creativity index of LLMs (GPT, Llama 2, OLMo) by an average of 30.1% at the verbatim level and 8.9% at the semantic levelGPT-3 / GPT base, Llama 2 / Llama 2 base, OLMo / OLMo base
- IC-418Matched n-grams in LLM outputs are concentrated in fewer reference documents than in human texts, indicating LLMs draw from a narrower set of sourcesGPT-3 / GPT base, Llama 2 / Llama 2 base, Tulu 2, OLMo / OLMo base
- IC-419CLIP ViT-B/16's layer-11 residual stream contains class-discriminative information in sparse SAE latent directions, and ablating class-specific top-k latents significantly degrades zero-shot classification accuracyCLIP / CLIP-ViT (LC)
- IC-420CLIP ViT-B/16's SAE latent interpretability is depth-dependent: layer 11 encodes semantic object concepts while layers 2, 5, and 8 encode local shapes and attention patternsCLIP / CLIP-ViT (LC)
- IC-421Sequential context-switching queries jailbreak Llama and Mistral models at 95% attack success rateLlama 3.1, Llama-3.2-3B, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 2 / Llama 2 base, Vicuna, Cohere Command R
- IC-422Safety fine-tuning in Llama models improves with parameter size but exhibits diminishing returnsLlama-3.2-3B, Llama 3.1
- IC-423Llama-3.1-8B-Instruct achieves 97.8% accuracy as a zero-shot toxicity classifier on ToxiGen, outperforming Llama-3-Guard-1B and matching Llama-3-Guard-8BLlama 3.1, Llama 3
- IC-424Structural in-context learning is transient in MultiBERTs and Pythia-1.4B, disappearing after early trainingMultiBERTs, Pythia
- IC-425Pretrained GPT-2 Large fails at structural in-context learning on unseen tokens in a syllogism taskGPT-2
- IC-426MultiBERTs exhibit a pushdown phenomenon where syntactic information migrates from later to earlier layers as training progressesMultiBERTs
- IC-427Newer base models (post-November 2023) outperform older ones by 7.3 points on MMLU and 19.1 points on GSM8K controlling for pretraining compute, but this gap vanishes after fine-tuning all models on the same task-relevant dataPythia, Llama 2 / Llama 2 base, Llama 3, Qwen1.5, Gemma, OLMo / OLMo base, StableLM, Falcon, GPT-J, InternLM, OpenLLaMA, RedPajama, Baichuan, Skywork, Yi, ZiYA2, MAP-NEO, Qwen 2
- IC-428Qwen 1.5 appears to Pareto-dominate Pythia and LLaMA 2 on MMLU and GSM8K, but after adjusting for test task training all three model families exhibit equivalent scalingPythia, Llama 2 / Llama 2 base, Qwen1.5
- IC-429The point of emergence for MMLU shifts from approximately 1.3×10²² flops to 5.6×10²⁰ flops as models train on 64,000 task-relevant examples, and the log-linear fit R² improves from 0.632 to 0.950Pythia
- IC-430Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signalGPT-2, Llama 2 / Llama 2 base, Llama 3.1
- IC-431Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD datasetLlama 2 / Llama 2 base
- IC-43256 LLMs from 19 families exhibit u-shaped scaling on hard questions and inverted-U scaling on easy questions, with the opposing trends explaining emergent ability stagnationGemma, Llama 2 / Llama 2 base, RedPajama-INCITE, Yi, StableLM, MPT, Falcon, Pythia, Qwen, Qwen1.5, BLOOM, DeepSeekMoE, OPT, GPT-Neo, CodeGen, XGLM, OpenLLaMA
- IC-433LLMs exhibit a non-monotonic ID-OOD performance gap (generalization valley) that peaks at intermediate task complexityQwen1.5, Llama-3.2-3B, Llama 3, Gemma 2, Claude 3, GPT-4o, O1 / OpenAI-o1-preview, Llama 3.1, Qwen2.5
- IC-434The critical complexity at which LLMs over-rely on memorization shifts to higher task difficulty as model size increasesQwen1.5, Llama 3.1, Gemma 2, Claude 3, GPT-4o, O1 / OpenAI-o1-preview
- IC-435Mistral-7B employs a less efficient algorithmic strategy (O(n²)) than Llama-3-8B (O([n², n³])) on probe tasks with multiple solution complexitiesMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3
- IC-436API selection accuracy of 10 LLM-based agents degrades sharply as task complexity increases, with open-source models ≥70B matching closed-source on simpler tasks but lagging on the most complexGemini 1.5 / Gemini Pro 1.5, Llama 3, Qwen 2, Qwen2.5, DeepSeek-2-Chat, DeepSeek-2-Coder, GPT-4o, GPT-3.5 / ChatGPT-3.5, GLM-4
- IC-437Extracting parameters from user queries is harder for LLM-based agents than using outputs from previous actions, and less intelligent LLMs show steeper parameter-filling degradation with task difficultyGemini 1.5 / Gemini Pro 1.5, Llama 3, Qwen 2, Qwen2.5, DeepSeek-2-Chat, DeepSeek-2-Coder, GPT-4o, GPT-3.5 / ChatGPT-3.5, GLM-4
- IC-438All 10 LLM-based agents perform poorly at recognizing when they need to request input from the system or user, with overall accuracy between 30.55% and 55.18%Gemini 1.5 / Gemini Pro 1.5, Llama 3, Qwen 2, Qwen2.5, DeepSeek-2-Chat, DeepSeek-2-Coder, GPT-4o, GPT-3.5 / ChatGPT-3.5, GLM-4
- IC-439Agent-specialized fine-tuned models (XLAM) significantly improve API selection over base models, but code-fine-tuned models (AgentLM) degrade performance, and no fine-tuning approach improves input recognitionAgentlm, Xlam-R, Lemur-v1-70B / Lemur-70B-Chat-V1, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-440GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet show up to 25% skill-level accuracy gaps despite overall accuracies within 0.4% of each otherGPT-4o, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5
- IC-441Skill-level improvements between model releases are highly uneven, with Claude 3.5 Sonnet gaining ~50% over Claude 3 Opus on law skills while Gemini improved most in math and scienceGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.5 / Gemini Pro 1.5, Gemini 1.0 Pro, Claude 3.5, Claude 3
- IC-442Routing each evaluation instance to the model strongest on its relevant skills yields a 3.2% accuracy gain over the best single model, with 3.5-6.8% gains on MMLU ProGPT-4o, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5
- IC-443Model inconsistency on probing questions negatively correlates with skill-slice accuracy (r = -0.675), with models contradicting themselves more often on skills where they perform poorlyGPT-4o, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5
- IC-444ViV1T produces non-differentiable population response representations when simulating mouse V1 experimentsViV1T
- IC-445Stable Diffusion v1.5 generates nudity for 796 out of 4703 prompts in the I2P inappropriate prompts datasetStable Diffusion
- IC-446In Stable Diffusion v1.5, concept-generating neurons are localized in the second layer of FFNs, spanning less than 3% of FFN parameters, and are disentangled from object-generating neuronsStable Diffusion
- IC-447CLIP ViT-B/16 produces noisy saliency maps and contains only 42 concept detectors, indicating poor visual interpretabilityCLIP / CLIP-ViT (LC)
- IC-448CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasetsCLIP / CLIP-ViT (LC)
- IC-449CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255CLIP / CLIP-ViT (LC)
- IC-450LLaVA with the original CLIP encoder produces noisy, non-sparse attention maps that poorly localize to the objects described in generated textLLaVA
- IC-451Transformer block coupling of Jacobian singular vectors positively correlates with benchmark performance across 30+ LLMs, more strongly than parameter count, depth, or embedding dimensionLlama 3, Llama 2 / Llama 2 base, Pythia, GPT-2, Gemma, MPT, Falcon, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Phi-2
- IC-452Transformer block coupling is absent at initialization and increases persistently throughout training in Pythia 12B and 6.9B, with layer-wise locality emergingPythia
- IC-453Hidden representation trajectories in trained LLMs exhibit considerable linearity (mean LSS 4.25) compared to 6.54 at initialization, and linearity increases with trainingLlama 3, Llama 2 / Llama 2 base, Pythia, GPT-2, Gemma, MPT, Phi-2
- IC-454Most hidden trajectories in trained LLMs exhibit exponential growth in norm as a function of depth, a property that emerges with trainingLlama 3, Llama 2 / Llama 2 base, Pythia, GPT-2, Gemma, MPT, Phi-2
- IC-455Qwen-Audio 7B zero-shot underperforms a 128M parameter baseline on audio difference explanation across three evaluation scenariosQwen-Audio
- IC-456VLM decoders achieve near-random accuracy on VALSE image-sentence alignment while pairwise accuracy is much higher, indicating reliance on linguistic priorsBakLLaVA, LLaVA-NeXT / LLaVA 1.6, mPLUG-Owl3
- IC-457All four tested VLM decoders are heavily text-centric when generating answers, with text modality contributing 85-97% of the prediction signalBakLLaVA, LLaVA-NeXT / LLaVA 1.6, mPLUG-Owl3
- IC-458Most VLM decoders show negative CC-SHAP on VALSE multiple-choice, indicating their explanations are less self-consistent than their answers, driven by a shift from text-dominant to image-dominant processingBakLLaVA, LLaVA-NeXT / LLaVA 1.6, mPLUG-Owl3
- IC-459Qwen-2.5 models show that the privacy-utility tradeoff for differentially private steering improves with model sizeQwen2.5
- IC-460Non-private activation steering of Llama-2-7B and Qwen-2.5-7B leaks membership information from the alignment dataset, while PSA reduces the empirical privacy lossLlama 2 / Llama 2 base, Qwen2.5
- IC-461Adding calibrated Gaussian noise to steering vectors (PSA) preserves alignment performance comparable to non-private mean steering across Llama-2-7B, Mistral-7B, Gemma-2-2B, and Qwen-2.5-7BLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma, Qwen2.5
- IC-462GPT-2 encodes toxicity in a low-dimensional linear subspace of its MLP layers, concentrated in higher layersGPT-2
- IC-463DPO's first-step gradients in GPT-2 are correlated with the toxic subspace, with stronger alignment in later layers and with more samplesGPT-2
- IC-464Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate methodVideo-LLaVA, Llama-VID
- IC-465GPT-4o achieves the highest weighted Cohen's kappa agreement with a reference-guided agent-debate method among all VLM judges, with scores exceeding 50 in several visual dimensionsVideo-LLaVA, Llama-VID, GPT-4o, InternVL2
- IC-466GPT-4o's evaluation reliability degrades when used as the final judge in a collective thought pipeline that aggregates reviews from less reliable VLMsLlama-VID, Video-ChatGPT, Video-LLaVA, GPT-4o
- IC-467Llama-3.1-405B's standard speculative decoding verification rejects correct continuations from GPT-4o, Llama-3.1-8B, and human text, accepting only roughly two tokens before the first rejection for GPT-4oLlama 3.1, GPT-4o
- IC-468Llama-3.1-405B's last hidden layer embeddings of erroneous tokens contain a linearly detectable error signal that a simple logistic regression head can exploit to flag incorrect continuationsLlama 3.1
- IC-469CLIP's global contrastive alignment causes attention on anatomically irrelevant regions in 3D CT, yielding limited zero-shot diagnostic accuracy (AUC 68.4 on 54 tasks)CLIP / CLIP-ViT (LC)
- IC-470LOVT and MGCA, which use implicit cross-attention local alignment, show only marginal improvement over CLIP in 3D CT diagnosis (AUC 69.4 and 70.1 vs 68.4)LOVT, MGCA, CLIP / CLIP-ViT (LC)
- IC-471Pythia models exceeding 100M parameters show a consistent leftward shift of the multifractal spectrum (increasing regularity) during training that is absent in the 14M and 31M variantsPythia
- IC-472The degree of emergence metric derived from Pythia's internal structure positively correlates with benchmark performance across training epochsPythia
- IC-473ResNet-18 lacks a clear multifractal structure while ResNet-152 shows one with irregular shifts, and a 160M diffusion model exhibits lower degree of emergence than Pythia 160MResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, Stable Diffusion, Pythia
- IC-474GPT-4, GPT-4o, and Llama-3.1-405B fail at knowledge classification and comparison without chain-of-thoughtGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Llama 3.1
- IC-475GPT-4, GPT-3.5, GPT-4o, and Llama-3.1-405B fail at inverse knowledge search regardless of promptingGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, GPT-4o, Llama 3.1
- IC-476GPT-4 and GPT-3.5 show strong positional bias in Chinese idiom character completionGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-477Released LLMs achieve limited success rates as web agents on WebArena-Lite, with open-source models substantially below proprietary onesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Llama 3.1, GLM-4, AutoWebGLM
- IC-478GPT-4 and GPT-4V achieve approximately 71-73% accuracy in judging whether a web agent trajectory successfully completes a taskGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-479SAM's mask decoder exhibits attention drift to background or specific object parts under imprecise prompts, causing severe segmentation degradationSAM, SAM 2
- IC-480Pre-trained M-LLMs (GPT-4o, LLaVA-v1.6-34B, InternVL2-26B, Qwen2-VL-7B) produce imprecise tampering explanations when artifacts require fine-grained pixel-level analysis such as lighting or perspective inconsistenciesGPT-4o, LLaVA-NeXT / LLaVA 1.6, InternVL2, Qwen2-VL
- IC-481Llama-3.1-8B and four other released LLMs reorganize their internal representations to reflect in-context graph structure in a sudden two-phase transition as context length increasesLlama 3.1, Llama-3.2-3B, Gemma 2
- IC-482In Llama-3.1-8B, in-context graph structure is absent in early layers dominated by semantic priors and emerges clearly in deeper layersLlama 3.1
- IC-483When in-context graph structure conflicts with pretrained semantic priors, Llama-3.1-8B encodes the in-context structure in higher principal components while the semantic prior dominates the first twoLlama 3.1
- IC-484Re-scaling the first 2-3 principal components of Llama-3.1-8B token representations causally shifts next-token predictions toward the target graph positionLlama 3.1
- IC-485LLMs show a significant performance gap between Wikipedia-based factual multi-hop QA and counterfactual multi-hop QA, indicating reliance on memorized knowledge rather than reasoning from contextGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Gemini, GPT-3 / GPT base, O1 / OpenAI-o1-preview, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Qwen 2
- IC-486LLMs achieve correct final answers through incorrect reasoning chains, inflating their apparent multi-step reasoning performanceGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Gemini, GPT-3 / GPT base, O1 / OpenAI-o1-preview
- IC-487Including sub-questions in the prompt improves LLM performance on multi-hop QA tasksGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Gemini, GPT-3 / GPT base, O1 / OpenAI-o1-preview
- IC-488LLM performance degrades progressively as the number of reasoning hops increases, with error propagation from earlier sub-questionsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Gemini, GPT-3 / GPT base
- IC-489State-of-the-art MLLMs fail at multi-step visual analogical reasoning, with best accuracy at 13% (Llama 3.2) on VOILA-WD and 29% (GPT-4o) on VOILA-ND, far below human performance of 71% and 70%GPT-4o, Llama 3.2, Qwen2-VL, CogVLM2, Seed-LLaMA-8B, MolmoE-7B, Emu2
- IC-490GPT-4o can identify visual relationships at 97% accuracy when given ground-truth descriptions but drops to 17% when asked to apply known relationships to new visuals, revealing a specific bottleneck in relational transferGPT-4o
- IC-491Presenting three images as a single collage rather than sequentially reduces MLLM accuracy by approximately 40% on the relationship application stepGPT-4o, Qwen2-VL, LLaVA-OneVision
- IC-492Least-to-most prompting consistently improves MLLM accuracy on the relationship application step compared to direct answering, with GPT-4o improving from 0.9% to 6.44% on VOILA-WDGPT-4o, CogVLM2, Seed-LLaMA-8B
- IC-493A linear direction in the input embedding space of Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-Instruct-v0.3, and Phi-3-mini-128k predicts instruction-following success, generalizes across tasks but not instruction types, and can be used to improve adherence via representation engineeringLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Phi-3
- IC-494The instruction-following dimension in Llama-2-7B-Chat and Llama-2-13B-Chat is more closely aligned with prompt phrasing than with task familiarity or instruction difficultyLlama 2 / Llama 2 base
- IC-495All evaluated multimodal foundation models achieve average non-hallucination accuracy below 50% across six hallucination scenariosFLUX / FLUX1, DALL·E 3, DALL·E 2, Stable Diffusion, Nova Pro, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama-3-2-Vision, LLaVA-NeXT / LLaVA 1.6, Gemini 1.5 / Gemini Pro 1.5
- IC-496GPT-4o achieves the highest location inference accuracy among evaluated models, reaching 98.16% for country, 60.23% for city, and 27.13% for zip code from street view imagesGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama-3-2-Vision, Nova Lite, Gemini 1.5 / Gemini Pro 1.5
- IC-497Text-to-image models experience performance drops exceeding 10% under adversarial prompts, with spatial reasoning being the most vulnerable task across all modelsNova Canvas, FLUX / FLUX1, DALL·E 3, DALL·E 2, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Stable Diffusion, LLaVA-NeXT / LLaVA 1.6
- IC-498Multimodal foundation models exhibit severe group unfairness, with race and age biases more pronounced than gender bias in text-to-image models while gender bias is stronger in image-to-text modelsFLUX / FLUX1, DALL·E 3, DALL·E 2, Stable Diffusion, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini 1.5 / Gemini Pro 1.5, Llama-3-2-Vision, Nova Canvas
- IC-499ViT patch embeddings contain local semantic information beyond the [cls] token, as shown by performance degradation when restricting the output head to [cls] only or removing positional embeddingsCLIP / CLIP-ViT (LC), ViT, MAE, DINOv2, SigLIP
- IC-500GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B solve multi-step retrieval problems without fine-tuning, achieving near-perfect accuracy for chains of up to 5 stepsGPT-4o, Claude 3.5, Gemini 1.5 / Gemini Pro 1.5, Llama 3.1
- IC-501Linear probes on middle-layer attention heads of Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 predict US lawmakers' DW-Nominate ideology scores with Spearman correlations around 0.85Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Vicuna
- IC-502Linear probes trained on US lawmaker ideology generalize to predict Ad Fontes media slant scores when the same models simulate news outletsLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Vicuna
- IC-503Adding probe regression coefficients to attention head activations steers Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 toward more liberal or conservative generated textLlama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Vicuna
- IC-504GPT-4o (gpt-4o-2024-08-06) rates the political slant of LLM-generated essays in close agreement with politically balanced human annotatorsGPT-4o
- IC-505Adversarial attacks on Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT shift hidden representations along the negative refusal feature directionLlama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma
- IC-506Restoring the refusal feature in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT causally disables all four tested adversarial attacksLlama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma
- IC-507The refusal feature direction in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT ranks near the top among 100 perturbations for compromising model safetyLlama 3, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Gemma
- IC-508GPT-4 exhibits reduced preference consistency (0.66 vs 0.84) when the quality distinction between two responses is minimalGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-509GPT-4 used as a preference labeler via prompt engineering yields alignment performance comparable to a task-specific 125M modelGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-510Qwen2-7B and Llama3-8B score near-random on textual temporal reasoning tasks while Qwen2-72B, Llama3-70B, and GPT-4o achieve near-perfect accuracy, showing temporal reasoning in LLMs is scale-dependent and emerges only above ~70B parametersQwen 2, Llama 3, GPT-4o, LongVA-7B, ViLA-8B
- IC-511LLaMA-2, Gemma, and Mistral all perform in-context density estimation via an adaptive kernel-like process, as revealed by their similar low-dimensional INPCA trajectories bounded between the geodesic and the Gaussian submanifoldLlama 2 / Llama 2 base, Gemma, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-518GPT-4o achieves 62.54% overall accuracy on MMWorld, the best among 15 MLLMs, while four open-source models perform below the 26.31% random-choice baselineGPT-4o, Claude 3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini, Video-LLaVA, Video-Chat-7B, Chat-UniVi-7B, mPLUG-Owl, Video-ChatGPT, PandaGPT-7B, ImageBind-LLM-7B, X-InstructBLIP-7B, LWM-1M-JAX, Otter-7B, Video-LLaMA-2-13B
- IC-519MLLMs exhibit different skill sets than humans, correctly answering expert-level questions that all three human annotators miss while failing on easy questions humans answer correctlyGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Claude 3.5, Gemini, Video-LLaVA, Video-Chat-7B, Video-ChatGPT, ImageBind-LLM-7B, PandaGPT-7B, Chat-UniVi-7B, Video-LLaMA-2-13B, X-InstructBLIP-7B, LWM-1M-JAX, Otter-7B, mPLUG-Owl
- IC-520Temporal reasoning performance drops significantly across all 15 MLLMs when video frames are shuffled or reduced to one-fifth of the original countGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3.5, Gemini, Video-LLaVA, Video-Chat-7B, Video-ChatGPT, ImageBind-LLM-7B, PandaGPT-7B, Chat-UniVi-7B, Video-LLaMA-2-13B, X-InstructBLIP-7B, LWM-1M-JAX, Otter-7B, mPLUG-Owl
- IC-521MLLMs show asymmetric modality-specific perception, with Gemini Pro achieving 69.97% on visual-only questions but only 24.45% on audio-only, while Video-Chat outperforms ChatUniVi on audio despite worse visual scoresGemini, Video-Chat-7B, Chat-UniVi-7B, Video-LLaMA-2-13B, Otter-7B
- IC-522LLMs perform correct example inference without inducing the correct rule, and this gap is robust to prompting methods, fact count, and scenario formGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o, Claude 3.5, Llama 3, Llama 2 / Llama 2 base
- IC-523LLMs rely on observed facts close to the test case in input feature space (neighbor-based reasoning) rather than on an abstract rule, and this effect is localizedGPT-4o, Claude 3.5, Llama 3, Llama 2 / Llama 2 base
- IC-525GPT-2 small's residual stream at layer 8 decomposes into two sub-spaces of approximately 25% and 75% of the dimensionalityGPT-2
- IC-526GPT-2 small's first token position has residual stream norms more than an order of magnitude larger than all other positionsGPT-2
- IC-527A 16 million latent sparse autoencoder substituted into GPT-4 yields a language modeling loss corresponding to 10% of GPT-4's pretraining computeGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-528The knowledge localization assumption fails for a large fraction of facts in GPT-2, Llama2-7B, and Llama3-8B, with 77% of facts classified as inconsistent knowledge in Llama3-8BGPT-2, Llama 2 / Llama 2 base, Llama 3
- IC-529For inconsistent knowledge in GPT-2, Llama2-7B, and Llama3-8B, the knowledge neurons are associated with the specific query rather than the fact, as shown by differential effects of suppressing or enhancing query-specific versus neighbor neuronsGPT-2, Llama 2 / Llama 2 base, Llama 3
- IC-530The attention module in GPT-2, Llama2-7B, and Llama3-8B plays a selective role in knowledge expression by activating specific knowledge neurons for a given query, as demonstrated by suppressing or enhancing attention scores at knowledge synapse positionsGPT-2, Llama 2 / Llama 2 base, Llama 3
- IC-531CONCH's zero-shot encoders cannot discriminate survival risk, achieving near-random concordance index on pathology whole-slide imagesCONCH
- IC-532PLIP's zero-shot encoders produce random-guessing-level survival predictions and consistently underperform CONCH on pathology survival analysisPLIP, CONCH
- IC-537Moirai's architectural enhancements (any-variate attention, multi-scale patch embedding, diverse mixture distribution) improve in-distribution forecasting but reduce out-of-distribution scalability relative to a simpler encoder-only baselineMoirai
- IC-538Chronos-T5's discrete probability prediction approach yields very small power-law exponents on NLL, limiting its scalability, and its in-distribution gains do not extend to out-of-distribution dataChronos
- IC-539GPT-4o in text-code-image mode achieves the highest scores on SCIMAGE but remains below 4 on all three evaluation dimensions, and all models degrade substantially on prompts requiring combined understanding typesGPT-4o, Llama 3.1, AutoTikZ / DataTikZ, DALL-E, Stable Diffusion
- IC-540Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image modelsGPT-4o, Llama 3.1, AutoTikZ / DataTikZ, Stable Diffusion, DALL-E
- IC-541Code-based output (Python or TikZ) produces images with notably higher scientific style scores than direct image generation from Stable Diffusion and DALL-EGPT-4o, Llama 3.1, Stable Diffusion, DALL-E
- IC-542The PyTorch pretrained ResNet50 on ImageNet is vulnerable to (1, y)-ACE calibration attacks that increase ECE from 3.70% to 47.23% while preserving accuracyResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-547Llama3.2 3B and Llama3.1 8B exhibit saturation events in which the top prediction, once it appears at a given layer, remains unchanged through all subsequent layersLlama-3.2-3B, Llama 3.1
- IC-548CLIP-B/32 exhibits progressively increasing layer-wise representation similarity in both its vision encoder and text encoder, and the pattern also holds across modalitiesCLIP / CLIP-ViT (LC)
- IC-549All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBenchGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, O1 / OpenAI-o1-preview, Claude 3.5, Llama 3.1, Llama 2 / Llama 2 base, Vicuna, Wizardlm, Qwen 2, Qwen1.5, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Phi-3, GLM-4, InternLM-2.5-7B
- IC-550Workflow generation performance scales with model size within families, but recently released 7B models outperform older 13B modelsQwen 2, Llama 3.1, Llama 2 / Llama 2 base, Wizardlm, Vicuna, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, InternLM-2.5-7B
- IC-551GPT-4's workflow generation performance declines as the number of nodes and edges in the workflow increasesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-552GPT-4, Llama-3.1-8B, and Qwen-2-72B all improve on ALFWorld and WebShop when given a generated workflow as structured prior knowledgeGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 3.1, Qwen 2
- IC-553OpenCLIP ViT-B/16's image-text alignment score is a strong predictor of domain generalization accuracy, while perceptual similarity to LAION-400M pre-training data is a weaker predictorOpenCLIP
- IC-554In Stable Diffusion 1.4, 1.5, 2.0, and 3.0, parameters with the smallest absolute values (below ~10^-3) do not contribute to the generative process, and this ineffectiveness is caused by stochastic training dynamics rather than architectural redundancy.Stable Diffusion
- IC-555Large LLMs (GPT-3.5-turbo, Gemini 1.5 Flash, Llama3-70B, Mixtral 46.7B) exhibit reasoning errors and significant accuracy degradation on large-scale logical commonsense reasoning tasks with 32k+ rules, even when the knowledge base is complete and retrieval is idealGPT-3.5 / ChatGPT-3.5, Gemini 1.5 / Gemini Pro 1.5, Llama 3, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, VERA
- IC-556All 9 LLM-based guard models exhibit significant miscalibration with average ECE exceeding 10% across 12 public benchmarks for both prompt and response classificationLlama Guard, Llama Guard 2, Llama-Guard 3, Aegis-Guard-Defensive, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, WildGuard
- IC-557Guard models show significantly degraded calibration under jailbreak attacks, with prompt classification ECE substantially higher than response classification ECELlama Guard, Llama Guard 2, Llama-Guard 3, Aegis-Guard-Defensive, WildGuard, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-558Guard models exhibit inconsistent calibration when classifying responses from different response model types, with ECE varying by up to 39 percentage points within a single modelLlama Guard, Llama Guard 2, Llama-Guard 3, Aegis-Guard-Defensive, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, WildGuard
- IC-559Contextual calibration is most effective for prompt classification while temperature scaling is more effective for response classification, but no single post-hoc method fully resolves miscalibrationLlama Guard, Llama Guard 2, Llama-Guard 3, Aegis-Guard-Defensive, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, WildGuard
- IC-560Llama3-8B-Instruct reliably distinguishes its own outputs from human outputs in self-recognition tasks, while Llama3-8B base performs at chance, indicating the ability is acquired during post-training.Llama 3
- IC-561A linear direction in the residual stream at layer 16 of Llama3-8B-Instruct is causally necessary and sufficient for self-authorship claims: steering with it achieves 100% control over authorship assertions, and projecting it out reduces claims by 50-60%.Llama 3
- IC-562Applying the layer-16 self-recognition vector to input tokens (not output) of Llama3-8B-Instruct alters the model's perception of authorship, making it believe or disbelieve it wrote arbitrary texts in both individual and paired paradigms.Llama 3
- IC-563The self-recognition vector's activation in Llama3-8B-Instruct is organized across depth: early layers (4-6) show diffuse perceptual activation to self-written text (present in both chat and base models), while layers 14-16 show a sharp decision-related peak at the output token that is present only in the chat model with role tags.Llama 3
- IC-564The implicit attention matrices of Mamba, RWKV, and Griffin exhibit depth-dependent structure, with dependencies between distant tokens becoming more apparent in deeper layersMamba, RWKV, Griffin
- IC-565Phi-3's residual stream encodes format instructions as linear directions, evidenced by cosine similarity and vocabulary-space projectionsPhi-3
- IC-566Adding instruction-specific steering vectors to the residual stream improves instruction-following accuracy for Phi-3, Gemma 2 2B IT, Mistral 7B IT, and Gemma 2 9B IT across format, length, and word-specific constraintsPhi-3, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-567Steering vectors computed on instruction-tuned Gemma 2 models transfer to base Gemma 2 models, with cross-model steering outperforming same-model steering for Gemma 2 2BGemma 2
- IC-568In Phi-3, word-exclusion steering vectors computed via difference-in-means project onto the vocabulary space with high logits for the excluded word, making them counterproductivePhi-3
- IC-569Mistral-7B-instruct-v0.1 achieves only F1 of 0.419 on zero-shot stance detection for the X-Stance German dataset, substantially below the fine-tuned BERT baseline (F1 0.693)Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-570Gradient-based image jailbreaks optimized against single or ensemble VLMs are universal for the attacked model(s) but do not transfer to other VLMs, except between highly similar modelsPrismatic, Qwen-VL, DeepSeek-VL
- IC-571Open-source VLMs (LLaVA, MiniGPT-4, InstructBLIP) are substantially more vulnerable to multimodal jailbreak attacks than Gemini-1.5-flash, with BAP attack ASR of 58–62% versus 40–41%LLaVA, MiniGPT-4, InstructBLIP, Gemini
- IC-572Bijection learning achieves state-of-the-art jailbreak ASR on frontier models, with peak ASR increasing with model capabilityClaude 3, Claude 3.5, GPT-4o
- IC-573Model capabilities on MMLU degrade monotonically as bijection encoding complexity increasesClaude 3, Claude 3.5, GPT-4o
- IC-574Guard models fail to effectively mitigate bijection attacks even at capability parity with the target modelClaude 3.5, GPT-4o, Claude 3
- IC-575Four released LLMs (LLaMA-3.1-8B, Mistral-7B, Qwen2-7B, Yi-1.5-9B) can perform in-context learning on continuous vector representations projected into their embedding space, matching or outperforming few-shot ICL across text, time-series, graph, and fMRI tasksLlama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Qwen 2, Yi
- IC-576For 10-digit numerical function regression, vector-ICL consistently outperforms few-shot ICL with raw number inputs across all four LLMs because continuous representations avoid multi-token splittingLlama 3.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Qwen 2, Yi
- IC-577An encoder's text reconstruction performance positively correlates with its effectiveness in downstream vector-ICL classification tasks across 15 encoder-LLM-dataset configurationsNV-Embed-v1, SFR-Embedding-2-R, Stella-en-1.5b-v5, GTR-T5-base
- IC-578LLMs encode input text as linearly separable representations in forerunner token hidden states, emerging in early layers and enhanced by in-context demonstrationsLlama 3, Falcon
- IC-579ICL hidden states exhibit positional bias: representations of the same input are more similar when the input appears at similar positions in the sequenceLlama 3
- IC-580The 3-step ICL inference circuit (text encoding, semantics merge, feature retrieval) is a dominant causal mechanism, as ablating the corresponding attention connections significantly degrades ICL accuracyLlama 3, Falcon
- IC-581Induction heads for ICL operate on task-specific attention subspaces, with partial overlap across tasks, and the geometry of these subspaces explains demonstration saturationLlama 3
- IC-582Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIsInstructBLIP, mPLUG-Owl, Idefics, ViT, CLIP / CLIP-ViT (LC), BLIP-2, Llama 2 / Llama 2 base
- IC-583Brain alignment in InstructBLIP and Idefics is organized by depth: middle layers align with higher visual regions while later layers align with early visual regions, whereas mPLUG-Owl shows later layers aligning with bothInstructBLIP, mPLUG-Owl, Idefics
- IC-584Most brain-explained variance is shared across task instructions, with image captioning (IC) acting as an umbrella category showing high overlap with VQ and CR but lower overlap with iu2 and srInstructBLIP
- IC-585MLLMs effectively capture count-related and recognition-related visual concepts with distinct brain alignment patterns, but produce similar alignment patterns for color, positional understanding, and general scene understandingInstructBLIP, mPLUG-Owl, Idefics
- IC-586Symbolic distance (number of reasoning steps) is the primary bottleneck for relational reasoning in LLMs, not total context lengthGemma, Mixtral, Gemini, GPT-4o
- IC-587Real-world knowledge acts as a shortcut in LLM relational reasoning, causing worse-than-chance performance on logically valid but factually incongruent statementsGemma, Mixtral, Gemini, GPT-4o
- IC-588Topologically ordered context improves relational reasoning over random ordering across nearly all LLMsGemma, Mixtral, Gemini, GPT-4o
- IC-589Flavor text (non-essential descriptive language) degrades relational reasoning in most LLMs, but GPT-4o is robust to itGemma, Mixtral, Gemini, GPT-4o
- IC-590Tuning only the identified safety neurons (SN-Tune) reduces harmful scores of instruction-tuned and base models by over 90 points while preserving general capability.Vicuna, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 2 / Llama 2 base, Llama 3
- IC-591Downstream fine-tuning on GSM8K degrades safety of Llama2-7b-chat and Mistral-7b-instruct-v0.2, but RSN-Tune partially preserves safety by protecting non-overlapping safety neurons.Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-592The log-likelihood layer in LLaMA-2-7B, LLaMA-2-7B-Chat, Vicuna-7B, and Mistral-7B-Instruct produces factually incorrect answers on TruthfulQA MC1 (817 samples) due to a misalignment between the output distribution and internal attention head representations, with LM-to-head-norm accuracy gaps of 24.23 to 40.68 points.Llama 2 / Llama 2 base, Vicuna, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Zephyr-7B-beta, Llama 3, Gemma 2
- IC-593The L2 norms of attention heads in Mistral-7B-Instruct and LLaMA-2-7B correlate with truthfulness, spiking by up to 83% at token positions of factual proposition completions and pertinent factual associations, and this correlation is specific to multi-headed attention representations rather than query, key, value, output, or FFN norms.Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 2 / Llama 2 base
- IC-594In LLaMA-2-7B, the truth-correlated attention heads are concentrated after layer 9, with two functional types (structural and associative) evenly distributed throughout the upper portions of the model, showing no further depth-dependent specialisation within that region.Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-595Gemini-pro has knowledge gaps on specific topics (Permian extinction, Fordism) causing it to perform far below its average rank on existing benchmarksGemini, Claude 2.0, GPT-3.5 / ChatGPT-3.5
- IC-596Multiple released LLMs fail to refuse harmful prompts disguised as historical or philosophical discussions, with GPT-4o and Mixtral showing the lowest refusal ratesGPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3, Claude 3
- IC-597Most LMMs exhibit systematic class bias in synthetic data detection, with GPT-4o biased toward classifying text as real and 3D as AI-generatedGPT-4o, Claude 3.5, Gemini 1.5 / Gemini Pro 1.5, InternVL2, Qwen2-VL, LLaVA-OneVision
- IC-598GPT-4o's synthetic image detection accuracy drops sharply on specialized domains (satellite 45.0%, medical 54.3%) compared to common image types (object 84.4%, person 84.4%)GPT-4o, Qwen2-VL
- IC-599All evaluated audio LMMs perform at or near random chance (44.4%–51.2%) on synthetic audio detection, while humans achieve 69.2%Qwen-Audio, SALMONN, OneLLM, Gemini 1.5 / Gemini Pro 1.5, AASIST
- IC-600Chain-of-thought prompting improves most LMMs on synthetic detection but degrades LLaVA-ov-7b from 56.6% to 18.8%, while GPT-4o performs well without it (64.1% baseline)GPT-4o, LLaVA-OneVision, InternVL2, Qwen2-VL, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5
- IC-601Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreasesQwen 2, Llama 3
- IC-602Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue growsQwen 2, Llama 3
- IC-603SIREN's embedding layer is not effectively optimized by gradient descent, so its embedding frequencies must be set as a hyperparameterSIREN
- IC-604CLIP ViT-B/32's CIFAR-10 image embeddings approximately satisfy a multi-cluster structure with near-orthogonal class-mean featuresCLIP / CLIP-ViT (LC)
- IC-605Six SOTA LLMs are vulnerable to composable jailbreak attacks, with maximum attack success rates ranging from 44% to 94%GPT-3.5 / ChatGPT-3.5, GPT-4o, Claude 3, Llama 3
- IC-606The relationship between model size and jailbreak vulnerability is reversed between Anthropic and Meta model familiesClaude 3, Llama 3, GPT-3.5 / ChatGPT-3.5, GPT-4o
- IC-607CLMBR-T-BASE's clinical prediction performance degrades as patient EHRs become more repetitive or irregularCLMBR-T-BASE
- IC-608Token trajectories in GPT-2, Llama 2 7B, Mistral 7B, and Llama 3.2 models cluster on a low-dimensional manifold and follow a linear drift plus Gaussian noise dynamicsGPT-2, Llama 2 / Llama 2 base, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3.2
- IC-609The last transformer layer of Mistral 7B, Llama 3.2 1B, and Llama 3.2 3B shows anomalous trajectory statistics inconsistent with the linear drift-plus-noise pattern of intermediate layersMistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3.2
- IC-610In LLaMA2-7B-Chat, RAG hallucinations are causally driven by copying heads losing external context information during generation and by knowledge FFNs in mid-to-upper layers over-adding parametric knowledge to the residual streamLlama 2 / Llama 2 base, Llama 3
- IC-620CAIT-S/24 (a ViT) attends more to high-frequency image features than ResNet-101 (a CNN), as revealed by frequency-domain attribution visualizationCAIT-S/24, ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-626Llama 7B retains 95% of its common sense reasoning performance when compressed to 2 GB using JLCMLLaMA
- IC-627ResNet 18 and ViT B/16 retain significant ImageNet accuracy at 2–3 bit weight compression via JLCMResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, ViT
- IC-632BLIP-2 succeeds on only 5 out of 100 advanced compositional vision-language tasksBLIP-2
- IC-633BLIP-2, LLaVA, and mPLUG-Owl show a trade-off between caption length and hallucination rate on COCOBLIP-2, LLaVA, mPLUG-Owl
- IC-634Video-ChatGPT achieves the highest correctness, detail, and contextual scores among four video understanding baselinesVideo-ChatGPT, VideoChat, LLaMA-Adapter, Video-LLaMA
- IC-635InstructBLIP achieves the highest overall MMBench score (44.0) among five evaluated vision-language modelsInstructBLIP, LLaVA, VisualGLM, OpenFlamingo
- IC-636Syntactic phenomena (determiner-noun and subject-verb agreement) localize to the same topmost-layer MLP neurons as factual information in BERT, GPT-2, and Llama-2BERT, GPT-2, Llama 2 / Llama 2 base
- IC-637KN edit (neuron suppression) has low reliability, overturning at most 5.2% of BLIMP categorical predictions and achieving only 1.66%–47.86% reliability on factual tasksBERT, T5, GPT-J
- IC-638ROME editing on GPT-2 XL and Llama-2 7B achieves high reliability but fails under bijective symmetry (23.71%–33.64%) and synonymous invariance (52.35%–58.36%) criteriaGPT-2, Llama 2 / Llama 2 base
- IC-639The causal tracing pattern of MLP at early layers and attention at late layers is not stable across factual and syntactic phenomena in GPT-2 XLGPT-2
- IC-642CLIP can infer contextual attributes (orientation, illumination, etc.) from images with approximately 74% accuracy on a binary taskCLIP / CLIP-ViT (LC)
- IC-643Conditioning CLIP on correct contextual attributes in the text prompt improves zero-shot classification accuracy across 13 image transformationsCLIP / CLIP-ViT (LC)
- IC-644CLIP relies on spurious features (background) as a shortcut in zero-shot classification, and conditioning on the correct background reduces this relianceCLIP / CLIP-ViT (LC)
- IC-645CLIP, PickScore, and HPSv2 text embeddings share a common direction (cone effect) that captures text-irrelevant preferences, and the orthogonal component c⊥p better measures T2I alignment; CLIP's untrained common direction makes it ineffective for reward fine-tuningCLIP / CLIP-ViT (LC), PickScore, HPSv2
- IC-646DINOv2, DeiT-III, and OpenCLIP repurpose approximately 2% of patch tokens in low-informative background areas as internal registers, discarding local patch information while aggregating global image information; DINO does not exhibit this behaviourDINOv2, DINO, DeiT-III, OpenCLIP, MAE
- IC-647DINOv2's feature-map artifacts cause it to be incompatible with the LOSt unsupervised object discovery method, scoring far below DINODINOv2, DINO, DeiT-III, OpenCLIP
- IC-662Translation ability in BLOOM models surges at approximately one-sixth of pre-training and then plateaus, with consistent dynamics across model sizes from 560M to 7.1BBLOOM
- IC-663OPT-1.3B, LLaMA-7B, and Aquila-7B encode sparse Harsanyi interactions, with only 29-51 salient interactions out of 1024 possible on SQuAD sentencesOPT, LLaMA, Aquila-7B
- IC-666DINOv2, CLIP-vision, and VGG-19 representations all align with MEG brain responses, with DINOv2 showing particularly high retrieval performance for late brain activity after image offsetDINOv2, CLIP / CLIP-ViT (LC), VGG / VGG13
- IC-667CLIP's intermediate layer features encode object boundaries recoverable by k-means clustering, a property absent in shallow and deep layersCLIP / CLIP-ViT (LC)
- IC-668SAM's edge-oriented segmentation yields high recall but very low precision because it cannot distinguish object boundaries from interior edgesSAM
- IC-669DINOv2's features, when clustered, produce smooth semantic regions but lack instance-level boundary delineationDINOv2
- IC-673Trained depthwise convolutional kernels in DS-CNN architectures converge to identifiable DoG-like patterns, with over 95% of ConvNeXtV2 and over 90% of ConvNeXt filters classifiable into a small set of clustersConvNeXtV2, ConvNeXt, MoGANet, ConvMixer, EfficientNet, MobileNetV3, MNASNet, ReplkNet-XL
- IC-677CLIP ViT's image representation is primarily constructed by the last 4 MSA layers, with MLPs and early MSA layers contributing negligiblyOpenCLIP
- IC-678Specific attention heads in CLIP ViT-L's last 4 layers encode specific image properties (color, shape, location, counting, texture) that are linearly recoverable via text directionsOpenCLIP
- IC-679CLIP relies on background/location as a spurious cue for bird classification, and ablating geolocation heads improves worst-group accuracy by 25.2%OpenCLIP
- IC-680CLIP ViT's image token contributions are spatially localized to match described content, enabling zero-shot segmentation that outperforms existing CLIP-based methodsOpenCLIP
- IC-681GPT-3.5 exhibits positional bias when judging which of two LLM responses is superiorGPT-3.5 / ChatGPT-3.5
- IC-682GPT-3.5 and GPT-4 achieve F1 scores of 0.5820 and 0.6180 respectively on pairwise response quality evaluation against human annotationsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-683LLaMA-7B, LLaMA-30B, Vicuna-7B, and Vicuna-13B achieve low accuracy (12.11% to 42.24%) in zero-shot and few-shot log-likelihood response evaluationLLaMA, Vicuna
- IC-684CLIP ViT-B/32 misclassifies 99% of forest satellite images as ocean when the word 'ocean' is overlaid as textCLIP / CLIP-ViT (LC)
- IC-685CLIP ViT-B/32 with a linear probe relies on gender as a spurious correlation for hair color, achieving only 15.85% accuracy on female gray hairCLIP / CLIP-ViT (LC)
- IC-686Adversarial perturbations alter CLIP ViT-B/32's token representations most strongly starting around layer 10CLIP / CLIP-ViT (LC)
- IC-687GPT-3.5+ models exhibit a gambler's fallacy bias and generate low-complexity sequences when asked to produce random binary sequencesGPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-688GPT-3.5-turbo-instruct-0914 shows sharp phase transitions in in-context learning of simple formal languages, transitioning from random generation to deterministic pattern repetition as context length increasesGPT-3.5 / ChatGPT-3.5, GPT-3 / GPT base
- IC-689Subjective randomness generation and sharp ICL transitions emerge only in larger or reward-fine-tuned models, absent in earlier GPT-3 variants and smaller open-source modelsGPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, text-ada-001, Llama 2 / Llama 2 base, Tulu 2, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b
- IC-698GPT-3.5-turbo's CoT reasoning errors are correlated across different demonstration sets, while PoT errors are less correlatedGPT-3.5 / ChatGPT-3.5
- IC-699Llama2-13b produces significantly less consistent answers than GPT-3.5-turbo on complex reasoning tasks, making it unsuitable as a weaker LLM in a cascadeLlama 2 / Llama 2 base, GPT-3.5 / ChatGPT-3.5
- IC-700GPT-4's reasoning accuracy degrades when provided with incorrect hints from a weaker modelGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-709Swin Transformer parameter redundancy is depth-dependent, with higher layers retaining fewer parameters than lower layers under module-aware pruningSwin Transformer
- IC-713Llama-2-7B-Instruct exhibits a reproducible failure mode in book-length summarization: high repetition and complete inability to perform incremental updatingLlama 2 / Llama 2 base
- IC-714For GPT-4 book-length summaries, human annotators prefer incremental summaries for detail (83% vs 11%) but hierarchical for structure (59% vs 35%), logic (53% vs 38%), and overall (54% vs 44%), showing coherence and human preference are not alignedGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-715Factual information deleted from GPT-J, LLaMA-2, and GPT-2-XL via ROME or MEMIT remains linearly recoverable from intermediate hidden states, with up to 89% extraction success at budget b=20GPT-J, Llama 2 / Llama 2 base, GPT-2
- IC-716Factual information deleted from GPT-J, LLaMA-2, and GPT-2-XL via ROME or MEMIT is recoverable by sampling outputs on automatically generated rephrased prompts, with up to 56% extraction success at budget b=20GPT-J, Llama 2 / Llama 2 base, GPT-2
- IC-717Llama-2-Chat's evaluation capability does not improve monotonically with model sizeLlama 2 / Llama 2 base
- IC-718GPT-4 achieves 0.882 Pearson correlation with human evaluators on 45 customized score rubrics while GPT-3.5-Turbo achieves only 0.392GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-719Llama-2-Chat achieves reasonable human-preference accuracy (51.78-53.67%) as a prompted reward model without specific reward-model trainingLlama 2 / Llama 2 base, StanfordNLP Reward Model, ALMOST Reward Model, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-721The 72-head entity tracking circuit identified in Llama-7b achieves high faithfulness in Vicuna-7b and Goat-7b without any modification to the circuit graphLLaMA, Vicuna, GOAT-7B
- IC-722Entity tracking in Llama-7b is implemented by detecting and transmitting the positional information of the correct entity, with distinct head groups for position detection, transmission, and value fetchingLLaMA, Vicuna, GOAT-7B
- IC-723The entity tracking performance gap between Goat-7b and Llama-7b is primarily attributable to enhanced positional information in the value fetcher and position transmitter headsLLaMA, GOAT-7B
- IC-730CLIPCap and BLIP-2 produce degraded alt-text on Twitter social media images, with BLEU@4 of 0.372 and 0.111 respectivelyCLIPCap, BLIP-2
- IC-731CLIP (ViT-B/32) achieves only 17.5 recall on video-text temporal alignment because it was trained on images and lacks video dynamicsCLIP / CLIP-ViT (LC)
- IC-732Off-the-shelf PyTorch ResNet classifiers are better calibrated than fine-tuned U-Net classifiers at high noise levels in the diffusion reverse processResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-733CoT prompting improves factual accuracy for instruction-tuned LLMs but degrades it for non-instruction-tuned LLMs such as OPT, BLOOM, and LLaMAOPT, BLOOM, LLaMA, Vicuna, ChatGLM-6B / ChatGLM-6b-2, FLAN-T5, GPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5
- IC-734GPT-3.5-turbo's factual verification F1 decreases as the number of reasoning hops required to validate a claim increasesGPT-3.5 / ChatGPT-3.5
- IC-735GPT-3.5-turbo's factual verification performance drops substantially under adversarial modifications, with man-made adversarial examples causing the largest declineGPT-3.5 / ChatGPT-3.5
- IC-736Vicuna-13B outperforms Vicuna-7B on factual knowledge tasks by 5.4% on averageVicuna
- IC-737CLIP ViT-B/16 binarized dot products yield 0.50–0.58 accuracy on binary concept presence queries across five image classification datasetsCLIP / CLIP-ViT (LC)
- IC-738BLIP-2 ViT-G FlanT5XL achieves 0.70–0.87 zero-shot accuracy on binary concept presence queries, competitive on most datasets but weaker on fine-grained CUB-200BLIP-2
- IC-739GPT-3.5-turbo-0613 combined with CLIP produces more faithful concept-salience pseudo-labels than LLaMA-2-13B-Chat, InstructBLIP, or LLaVA-1.5B on most of five datasetsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base, InstructBLIP, LLaVA, CLIP / CLIP-ViT (LC)
- IC-742ResNet-50-BN on Waterbirds relies on background as a spurious feature for classification, and this shortcut is invisible to entropy-based confidence metricsResNet / ResNet-152 / ResNet-101 / ResNet-50-BN
- IC-744The decoded vocabulary of a function vector often reflects the task's output space, but reconstructing a vector that matches this vocabulary distribution does not recover the FV's full causal effect.GPT-J
- IC-745Function vectors for simple list-oriented tasks can be algebraically combined via addition and subtraction to produce new vectors that trigger composed tasks, sometimes outperforming 10-shot ICL.GPT-J, Llama 2 / Llama 2 base
- IC-746RoBERTa-Large pretrained with different mask ratios exhibits a sweet spot in downstream accuracy on QNLI and SST-2RoBERTa / RoBERTa-L
- IC-747SOTA pruning methods (SparseGPT, Wanda, magnitude) cause significant degradation on knowledge-intensive tasks for Vicuna and Llama models at 25-30%+ unstructured sparsity, and fail completely for n:m structured sparsityVicuna, LLaMA, Llama 2 / Llama 2 base
- IC-748Pruned LLMs at ≥50% sparsity remain robust in-context retrievers and summarizers, with Vicuna-7B matching up to ~40% sparsity and Vicuna-13B up to ~50% sparsity in open-book settingsVicuna
- IC-749Compressed Vicuna-13B at 46.16% sparsity (matching 7B parameter count) achieves lower MMLU accuracy than dense Vicuna-7B, indicating large-sparse models do not outperform small-dense at matched sizeVicuna
- IC-750Open-source models without safety training are significantly more vulnerable to jailbreak attacks than safety-aligned proprietary modelsVicuna, Alpaca, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base
- IC-751Arena-Hard-200 reveals larger performance gaps between open and proprietary LLMs than MT-BenchGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Claude Instant 1, Vicuna, Llama 2 / Llama 2 base, Wizardlm
- IC-752GPT-4's win rate over GPT-3.5-turbo is 52% on the top-50 most challenging prompts but only 22% on the bottom-50 easiest promptsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-761GPT-3.5-turbo and GPT-4 are susceptible to specific circulating jailbreaking prompts, with 'jailmommy' achieving a 71.16% success rate in producing toxic outputsGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-762GPT-4 and GPT-3.5 outperform humans in generation but underperform in discriminative (selective) evaluation across 10 of 13 language tasksGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-763CLIP and OpenCLIP fall short of human discriminative accuracy on vision tasks, with performance dropping substantially under hard negativesCLIP / CLIP-ViT (LC), OpenCLIP
- IC-764GPT-4 and GPT-3.5 make frequent errors answering questions about their own generated text, underperforming humans in interrogative evaluationGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-765BLIP-2, BLIP, InstructBLIP, Bard, and BingChat fall short of human accuracy in answering questions about Midjourney-generated imagesBLIP-2, BLIP, InstructBLIP, Bard, BingChat
- IC-766Multiple Real-SR methods fail to outperform a small FSRCNN network on the majority of 100 representative degradation casesSRResNet, DASR, RD-SR, ESRGAN
- IC-767BSRNet outperforms RealESRNet on the majority of degradation cases, reversing the ranking obtained from a single random test setBSRNet, RealESRNet
- IC-768MMRealSR exhibits the most consistent performance across degradation cases while SwinIR achieves the highest performance on cases it handles well, among GAN-based Real-SR methodsMMRealSR, SwinIR
- IC-783Stable Diffusion v1.5's conditional probability pθ(x|c) is heavily biased by the unconditional probability pθ(x), making it unreliable as a condition-alignment metricStable Diffusion
- IC-784Pre-trained scoring models (CLIP Score, HPS, Image Reward, Pick Score) underperform on domain-specific fine-tuned diffusion modelsCLIP / CLIP-ViT (LC), HPS, ImageReward, PickScore, Van Gogh Diffusion
- IC-785Different released diffusion models produce images with distinguishable probability signatures, enabling source attributionDreamlike Photoreal 2.0, OpenJourney, Stable Diffusion
- IC-786CLIP-ViT (LC) achieves 0.87 accuracy and 0.91 average precision on fake image detectionCLIP / CLIP-ViT (LC)
- IC-788CLIP ViT-B/32 fails to retrieve the correct image even when the generated target caption is well-aligned with the ground-truth imageCLIP / CLIP-ViT (LC)
- IC-789CLIP retrieval quality in zero-shot compositional image retrieval scales log-linearly with model size from approximately 150M to 2.5B parametersCLIP / CLIP-ViT (LC), OpenCLIP
- IC-790GPT-4 outperforms GPT-3.5-turbo, Vicuna-13B, and Llama2-70B for generating target captions in zero-shot compositional image retrievalGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Vicuna, Llama 2 / Llama 2 base
- IC-791BLIP-2, BLIP, and COCA produce captions of comparable quality for zero-shot compositional image retrievalBLIP-2, BLIP, COCA
- IC-7951D subspaces of MLP activations found by DAS in GPT-2 Small (IOI) and GPT-2 XL (factual recall) produce apparent causal effects that are interpretability illusions driven by causally disconnected components activating dormant pathwaysGPT-2
- IC-796GPT-2 Small MLP weight matrices are full-rank across all 12 layers and residual stream features are linearly recoverable from post-GELU MLP hidden activations, providing the structural conditions for the subspace patching illusionGPT-2
- IC-798GPT-4 achieves 61.6% on MATH with tool-integrated reasoning prompting, exceeding PaL (51.8%) and CoT (42.5%)GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-799WizardMath-70b scores lower than base Llama-2-70b on TabMWP (49.8% vs 57.5%), indicating degraded OOD generalization from rationale-based fine-tuningWizardMath, Llama 2 / Llama 2 base
- IC-807Language-conditioned robot policies RT-1 and RT-2 fail to generalize to unseen manipulation tasks, achieving only 16.7% and 11.1% success rates respectively across 7 novel skillsRT-1, RT-2
- IC-808Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layersPythia
- IC-809Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergencePythia
- IC-810Individual sparse autoencoder features in Pythia-70m-deduped are monosemantic and have predictable causal effects on output logits, as demonstrated by an apostrophe feature whose ablation primarily suppresses the 's' tokenPythia
- IC-811Alpaca's 52k instruction-tuning data is predominantly low-quality (only 17.75% score ≥ 4.5 on accuracy), yet the full 52k data still yields higher MMLU scores than the filtered 9k subset for both 7b and 13b variantsAlpaca
- IC-812InstructGPT (text-davinci-003) reduces content diversity in co-written essays while GPT-3 (davinci) does not, and the effect is attributable to the model's own less diverse text contributionsInstructGPT, GPT-3 / GPT base
- IC-815RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B modelsPythia, LLaMA
- IC-816RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B modelsPythia, LLaMA
- IC-817RLHF on general-purpose preference data improves machine ethics performance in Pythia and Llama-7B modelsPythia, LLaMA
- IC-818RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B modelsPythia, LLaMA
- IC-824GPT-3 models of all sizes (350M to 175B) can reverse name-description associations in-context with near-perfect accuracy, showing the reversal curse is a property of training rather than reasoningGPT-3 / GPT base
- IC-827LLMs exhibit distinct psychological profiles that differ from human norms and vary by model size and versionGPT-3 / GPT base, GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base
- IC-828Jailbreaking GPT-4 via cipherchat shifts its psychological profile toward human norms and reduces emotional intelligence scoresGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-829Role assignment to GPT-3.5-turbo produces role-consistent changes in psychological profiles and task performance, validating the psychometric scalesGPT-3.5 / ChatGPT-3.5
- IC-834CLIP zero-shot predictions exhibit high equal opportunity difference when target and sensitive attributes are intrinsically dependentCLIP / CLIP-ViT (LC)
- IC-835CLIP zero-shot predictions exhibit large worst-group accuracy gaps due to spurious correlations on Waterbirds and CelebACLIP / CLIP-ViT (LC)
- IC-836CLIP zero-shot predictions exhibit demographic bias on FairFace when using attribute-unrelated text promptsCLIP / CLIP-ViT (LC)
- IC-837CLIP ViT-L/14 embeds more target-attribute information and less sensitive-attribute information than CLIP ResNet-50 on WaterbirdsCLIP / CLIP-ViT (LC)
- IC-840GPT-2 Small's name mover heads exhibit disrupted attention patterns under out-of-distribution Gaussian noise corruptionGPT-2
- IC-847CLIP ViT-B/32 CLIPScore achieves only ρ=0.276 / τ=0.191 correlation with human 1-5 likert T2I alignment ratings on TIFA160CLIP / CLIP-ViT (LC)
- IC-848GPT-3.5 achieves 98.3% precision and 96.0% recall for automatic question-tuple matching but makes errors when questions differ in wording yet are semantically uniqueGPT-3.5 / ChatGPT-3.5
- IC-849GPT-2 and T5-base exhibit vanishing expected gradients under RFT for inputs with small reward standard deviation, prevalent in 3 of 7 GRUE datasets, causing RFT to underperform SFTGPT-2, T5
- IC-850A partial SFT phase (40% of steps, 1% of samples) before RFT allows GPT-2 and T5-base to reach 96% of the reward achieved with full SFT+RFT, by reducing the number of inputs with vanishing gradientsGPT-2, T5
- IC-852Intrinsic self-correction without external feedback consistently degrades reasoning accuracy across GPT-3.5-turbo, GPT-4, GPT-4-turbo, and LLaMA-2-70B-chatGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base
- IC-853Multi-agent debate with GPT-3.5-turbo-0301 does not outperform self-consistency at equivalent inference cost on GSM8KGPT-3.5 / ChatGPT-3.5
- IC-854The apparent self-correction improvement in constrained generation (Madaan et al., 2023) is an artefact of a sub-optimal initial prompt, not a genuine model capabilityGPT-3.5 / ChatGPT-3.5
- IC-858The choice of graph encoding method significantly changes LLM accuracy on graph reasoning tasks, with incident encoding outperforming adjacency by up to 34 percentage points on connected nodesPaLM 62B, PaLM 2, GPT-3.5 / ChatGPT-3.5
- IC-859LLMs rely on learned priors about graph properties (cycles exist, edges are absent) rather than analyzing the specific graph structure, causing below-majority-baseline performance and extreme structure-dependent accuracyPaLM 62B
- IC-860Larger PaLM 2 models (xxs to l) show progressively better graph reasoning, but even the largest variant fails to beat the majority baseline on edge existencePaLM 2
- IC-861LLMs achieve near-zero accuracy on the disconnected nodes task, indicating an inability to reason about the absence of edges in a graphPaLM 62B
- IC-866A prefix applied to Llama-7B's first attention layer preserves the relative attention distribution over content positions and only adds a constant-direction bias to the attention block outputLLaMA
- IC-867In GPT-2 prefix-tuned on the emotion dataset, attention over prefix positions is nearly constant across inputs, collapsing the effective bias subspace to a single direction in most layersGPT-2
- IC-868Skill-Mix performance degrades with increasing k, and within the Llama-2 family the saturation point increases with model sizeLlama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Falcon, Xwin-LM-70B-v0.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Qwen, TigerBot-70B-Chat
- IC-869Models ranking highly on popular LLM leaderboards perform worse than Llama-2-70b-chat on Skill-Mix, suggesting cramming for the leaderboard at the expense of general-purpose text generationFalcon, Xwin-LM-70B-v0.1, Qwen, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, TigerBot-70B-Chat, Llama 2 / Llama 2 base
- IC-870GPT-4's performance on Skill-Mix(k=5) and Skill-Mix(k=6) provides probabilistic evidence of generating novel skill-topic combinations not present in training dataGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-871Llama-2-70b-chat as a grader is more generous than GPT-4 and systematically gives higher scores to Llama-2 family outputsLlama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
- IC-879Code Llama outperforms Llama-2 on coding (HumanEval) and mathematical (GSM8K) reasoning at both 7B and 13B scalesLlama 2 / Llama 2 base, Code Llama
- IC-887HuggingGPT's in-context task-model assignment always selects the same model regardless of input question or task typeHuggingGPT
- IC-888SLIMG fails on heterophily graphs for link prediction because it cannot properly measure node similarity of heterophily embeddingsSLIMG
- IC-894FLAN-T5 base's average per-token probability increases with token index during generation on WMT translation tasksFLAN-T5
- IC-895Across all FLAN-T5 sizes, BLEURT scores show a negative correlation with prediction length on WMT translation tasksFLAN-T5
- IC-897GPT-4 underperforms codex on Spider text-to-SQL with few-shot prompting, attributed to its zero-shot tuningGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, code-davinci-002
- IC-898GPT-3.5-turbo and GPT-4 are overconfident in their initial code predictions when unit test execution is unavailableGPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-899GPT-4, when used as a blind pairwise evaluator, exhibits the same style-over-factuality preference as human crowdworkersGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-900BLIP-2, MiniGPT-4, and LLaVA-1.5 show degraded zero-shot VQA accuracy on underspecified questions, with absolute improvements of 1.14–7.94% when questions are augmented with visually-grounded detailsBLIP-2, MiniGPT-4, LLaVA-1.5 / LLaVA-v1.5
- IC-901BLIP-2's LLM-only VQA performance improves with more specified questions while the image remains essential, revealing asymmetric strength between the LLM and vision componentsBLIP-2
- IC-902BLIP-2 and MiniGPT-4 confidence-based question selection underperforms the original question for paraphrased candidates but succeeds for semantically enriched REPARe questionsBLIP-2, MiniGPT-4
- IC-903Knowledge editing performance (ES, GS, LS) improves as model scale increases from GPT-2 (124M) to T5-XL (2.8B) to GPT-J (6B) across all editing methodsGPT-2, T5, GPT-J
- IC-904Stable Diffusion XL generates non-empty cups when prompted for 'empty cup'Stable Diffusion
- IC-913In OPT-2.7B, Pythia-70M/1.4B/6.9B, and BERT-base, the stable rank of MLP lower layers shows a drop-and-bounce pattern during training that is more salient in top layers while bottom layers show suppressed dropping curvesOPT, Pythia, BERT
- IC-914In Pythia models (70M through 2.8B), BERT-base, OPT-6.7B, LLaMA-2-7B, and ViT-Huge, the MLP out-projection vectors are almost orthogonal throughout trainingPythia, BERT, OPT, Llama 2 / Llama 2 base, ViT
- IC-915In Pythia-70M and Pythia-160M, individual MLP hidden neurons are activated by multiple irrelevant token combinations (pattern superposition)Pythia
- IC-916Pythia models show scale-dependent last-layer averaging barriers: 70M exhibits a barrier of ~13 while 410M shows ~1Pythia
- IC-917ViT-S models show early-layer sensitivity to layer-wise averaging, with the averaging direction being far more disruptive than random perturbations of the same normViT
- IC-925MultiBERTs exhibits the same phase transition pattern (UAS spike, loss drop, BLIMP improvement) as the authors' own BERT-base trainingMultiBERTs
- IC-926Across 25 MultiBERTs seeds, UAS does not correlate with MLM test loss or BLIMP accuracyMultiBERTs
- IC-927Human ciphers (ASCII, Unicode, Caesar, Morse) bypass the safety alignment of GPT-4 and GPT-3.5-turbo, with more powerful models producing more unsafe responsesGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Falcon, Llama 2 / Llama 2 base, GPT-3 / GPT base
- IC-928SelfCipher (a role-play prompt without explicit cipher rules) evokes a 'secret cipher' in LLMs, achieving high unsafety rates that outperform most human ciphersGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Falcon, Llama 2 / Llama 2 base, GPT-3 / GPT base
- IC-929Simulated character-level ciphers that never appear in pretraining data cannot bypass safety alignment even with 10+ demonstrationsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
- IC-930StackLLama, when used as a reward model, achieves near-random consistency on contrast instructions for the Stack Exchange taskStackLLaMA
- IC-931GPT-4 achieves approximately 95% accuracy on contrast instructions, far exceeding human performance without toolsGPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
- IC-932Pre-trained DNN object detectors show a sharp falloff in peripheral detection performance with increasing eccentricity, degrading to near-chance by 20°, while human performance degrades graduallyDINO-FocalNet-Large, Swin Transformer, DETR-R50, RetinaNet-R50, FoveaBox, Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101
- IC-933DNN object detectors do not exhibit the same sensitivity to image clutter as humans in peripheral object detectionDINO-FocalNet-Large, Swin Transformer, DETR-R50, RetinaNet-R50, FoveaBox, Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101
- IC-934DNN object detectors do not exhibit the same object size effect as humans in peripheral object detectionDINO-FocalNet-Large, Swin Transformer, DETR-R50, RetinaNet-R50, FoveaBox, Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101
- IC-935In pre-trained ViT, query vector Kruskal rank reaches the context size only after one self-attention layer, while general position fails at all depthsViT
- IC-936GPT-2's learned positional encodings cause context vectors to lose linear independence after one layer, whereas BERT's sinusoidal encodings preserve itGPT-2, BERT
- IC-939CLIP reward landscapes are well-shaped for photorealistic environments but poorly shaped for abstract renderingsCLIP / CLIP-ViT (LC)
- IC-940CLIP can specify 5 of 8 complex humanoid tasks from single-sentence prompts, failing on tasks requiring discrimination of subtle body-pose differencesCLIP / CLIP-ViT (LC)
- IC-941CLIP reward model quality scales with model size, with a sharp phase transition between ViT-H/14 and ViT-BigG/14 for the humanoid kneeling taskCLIP / CLIP-ViT (LC)
- IC-945LLaMA-2, MPT, Falcon, Pythia, and BERT-base-uncased allocate disproportionate attention to initial tokens regardless of their semantic contentLlama 2 / Llama 2 base, MPT, Falcon, Pythia, BERT
- IC-946LLaMA-2-7B, MPT-7B, Falcon-7B, and Pythia-12B do not consistently improve in perplexity as the StreamingLLM cache size increasesLlama 2 / Llama 2 base, MPT, Falcon, Pythia
- IC-951The programmatic space achieves behavior-similarity values comparable to the LEAPS latent space, indicating that optimizing the behavior loss alone does not produce a more search-conducive spaceLEAPS
- IC-955All baseline knowledge tracing models either lack significant correlation or negatively predict causal support from their inferred prerequisite graphsHLR, DKTF, AKT, GKT, QIKT
- IC-959LLaMA and OPT-1.3B (and Aquila-7B) encode more similar interaction primitives than smaller models such as BERT-base and BERT-largeLLaMA, OPT, Aquila-7B
- IC-981ImageBind's indirect alignment through images degrades zero-shot performance on non-visual modalities and prevents emergent cross-modal retrievalImageBind
- IC-982In Stable Diffusion's UNET, visual attribute knowledge is distributed across multiple components with attribute-specific patterns, concentrated more in the up-block, and cross-attention layers are not the primary causal statesStable Diffusion
- IC-983In Stable Diffusion's CLIP text-encoder, knowledge about all visual attributes is localized to a single causal state: the first self-attention layer at the last subject tokenStable Diffusion, CLIP / CLIP-ViT (LC)
- IC-984CLIP ViT-B/16's representation space does not reliably preserve semantic similarity as measured by shared image tagsCLIP / CLIP-ViT (LC)
- IC-985LLaMA-2-7B plateaus in ICL accuracy and fails to override semantic priors when in-context labels are flipped on a simple happy/sad classification taskLlama 2 / Llama 2 base
- IC-986Most LLMs lack tool usage awareness, with only ChatGPT exceeding 70% F1 in zero-shot evaluationChatGPT, ChatGLM2, Llama 2 / Llama 2 base, Vicuna, Koala
- IC-987When the correct tool is absent from the candidate list, most LLMs hallucinate a tool rather than returning 'none'ChatGPT, ChatGLM2, Llama 2 / Llama 2 base, Vicuna, Koala
- IC-988LLMs show large gaps in multi-tool selection and over-rely on the number of tools specified in the promptChatGPT, ChatGLM2, Llama 2 / Llama 2 base, Vicuna, Koala
- IC-989Tool selection CSR degrades as the candidate tool list grows from 5 to 15 tools, and performance varies by user scenarioChatGPT, ChatGLM2, Llama 2 / Llama 2 base, Vicuna, Koala
- IC-990ResNet-50 and DenseNet-101 exhibit a higher mean-to-variance ratio in penultimate pre-ReLU activations for in-distribution samples than for out-of-distribution samples, and the activation-based scaling factor is well-separated between ID and OODResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, DenseNet / DenseNet-101
- IC-991LLaMA-2, Falcon-7B, and GPT-3.5-Turbo exhibit large performance spread (up to 76 accuracy points) across semantically equivalent prompt formats, and model comparison rankings are frequently reversed by format choiceLlama 2 / Llama 2 base, Falcon, GPT-3.5 / ChatGPT-3.5
- IC-992LLaMA-2-7B's last hidden layer encodes the prompt format with high identifiability, and the separability of format embeddings in the top two principal components correlates with performance spreadLlama 2 / Llama 2 base
- IC-996GPT-3.5 achieves 73.5% zero-shot accuracy on OGBN-ARXIV 40-class node classification and 73.56% on TAPe-ARXIV23, a dataset of papers published after its knowledge cutoffGPT-3.5 / ChatGPT-3.5
- IC-997LLaMA-2-13B-Chat achieves 44.23% zero-shot accuracy on OGBN-ARXIV, substantially below GPT-3.5's 73.5%Llama 2 / Llama 2 base
- IC-998GPT-3.5's zero-shot accuracy on OGBN-ARXIV depends on the position of the title relative to the abstract in the prompt: 0.720 when abstract precedes title, 0.695 when title precedes abstractGPT-3.5 / ChatGPT-3.5
- IC-999FLUX1 and Stable Diffusion 3.5 exhibit high local dependency ratio and produce text hallucinations when generating text contentFLUX / FLUX1, Stable Diffusion
- SY-001Objects outside the patient, gown snaps and ECG electrodes, drive Sybil's risk predictionsSybil
- SY-002Sybil responds more weakly to nodules near the pleura, where adenocarcinoma tends to appearSybil
- SY-003Sybil processes pulmonary nodules almost additively, with limited pairwise interactionsSybil
- TM-001A mean-difference vector between before and after image pairs acts as a transferable concept vectorTerraMind
- TM-002The mean-difference concept vector scores higher than trained classifiers under cosine similarityTerraMind
- TM-003Optical-to-SAR generation errors cluster spatially, but only five regions survive FDR controlTerraMind
- TM-004Surface composition, not the acquisition time gap, correlates with optical-to-SAR reconstruction errorTerraMind
- TM-005Flooded vegetation gives the lowest reconstruction error of any land-cover class, not the highestTerraMind
- TM-006Coordinates are recoverable from TerraMind's frozen features, latitude more accurately than longitudeTerraMind
- TM-007Coordinates are linearly decodable only in the larger TerraMind variantsTerraMind
- TM-008TerraMind embeddings separate hemispheres rather than climate zonesTerraMind
- TM-009A weak seasonal shift in the embeddings matches climatic intuition but is attributed to geographyTerraMind
- TM-010Fixed corner patches dominate Integrated Gradients maps regardless of image contentTerraMind
- TM-011Spatial coherence of latent feature planes increases with encoder depthTerraMind
- TM-012Longitude and time-of-year planes become circular in deeper encoder blocksTerraMind
- TM-013Intervening on the altitude plane raises generated terrain by almost 500 metresTerraMind
- TM-014Per-edge latent distance along shortest paths spikes at physical barriersTerraMind