Light Dark A finding is a claim about how one model behaves, made by someone other than its authors, after the fact. Findings about different models meet on the mechanisms they describe.
Models Llama 2 / Llama 2 base text 207 findings GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report text, image 179 findings GPT-4o text, image 136 findings GPT-3.5 / ChatGPT-3.5 text 133 findings Llama 3 text 121 findings Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 text 89 findings Llama 3.1 text 81 findings CLIP / CLIP-ViT (LC) image, text 80 findings Pythia text 64 findings GPT-2 text 63 findings Vicuna text 56 findings LLaMA 48 findings Claude 3.5 text, image 47 findings Claude 3 text, image 41 findings Gemma 2 text 41 findings Gemini 1.5 / Gemini Pro 1.5 text, image 38 findings Stable Diffusion image 37 findings GPT-3 / GPT base text 34 findings ViT image 34 findings Qwen 2 text 33 findings Gemma text 30 findings ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN image 29 findings GPT-J text 27 findings ChatGPT 26 findings Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b text 24 findings OPT text 24 findings Gemini text, image 22 findings Falcon text 21 findings LLaVA-1.5 / LLaVA-v1.5 text, image 19 findings BLIP-2 text, image 18 findings LLaVA-NeXT / LLaVA 1.6 text, image 18 findings DINOv2 image 17 findings InstructBLIP text, image 17 findings Phi-3 text 16 findings MPT text 15 findings O1 / OpenAI-o1-preview text 15 findings OpenCLIP text, image 15 findings Qwen2.5 text 15 findings OLMo / OLMo base text 14 findings Qwen1.5 text 14 findings TerraMind image 14 findings Mixtral text 13 findings PaLM 2 text 13 findings BERT text 12 findings Qwen2-VL text, image 12 findings InternVL2 text, image 11 findings LLaVA text, image 11 findings GLM-4 text 10 findings BLOOM text 9 findings VGG / VGG13 image 9 findings Yi text 9 findings mPLUG-Owl text, image 9 findings Qwen text 8 findings Swin Transformer image 8 findings CodeLlama-13B text 7 findings Idefics image, text 7 findings Llama 3.2 text 7 findings Llama-3.2-3B text 7 findings MAE image 7 findings MiniGPT-4 text, image 7 findings SAM image 7 findings Code Llama 6 findings ConvNeXt image 6 findings DINO image 6 findings Gemini 1.0 Pro text, image 6 findings MultiBERTs text 6 findings OpenFlamingo 6 findings Qwen-VL text, image 6 findings RoBERTa / RoBERTa-L text 6 findings SigLIP image, text 6 findings Tulu 2 text 6 findings Video-LLaVA text, video 6 findings WildGuard text 6 findings mPLUG-Owl3 text, image 6 findings Cohere Command R text 5 findings DeepSeek-VL text, image 5 findings EVA-CLIP text, image 5 findings InstructGPT 5 findings Llama Guard 2 text 5 findings Llama-Guard 3 text 5 findings Qwen 2.5 72B Instruct text 5 findings Video-ChatGPT text, video 5 findings Wizardlm text 5 findings Aegis-Guard-Defensive text 4 findings Alpaca 4 findings Chat-UniVi-7B text, image, video 4 findings ChatGLM2 4 findings Claude 2.0 text 4 findings CodeLlama-34B 4 findings CogVLM2 text, image 4 findings Command R+ text 4 findings DeepGate2 graph 4 findings ESM-2 text 4 findings EfficientNet image 4 findings FLAN-T5 4 findings FLUX / FLUX1 text, image 4 findings Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 image 4 findings Idefics2 text, image 4 findings InternLM-XComposer2-VL text, image 4 findings Koala 4 findings LLaVA-OneVision text, image 4 findings Llama Guard text 4 findings Llama-3-2-Vision text, image 4 findings MiniCPM-V text, image 4 findings Otter 4 findings Otter-7B text, image 4 findings SALMONN audio, text 4 findings T5 4 findings TSM video 4 findings Video-Chat-7B video, text 4 findings Video-LLaMA-2-13B video, audio, text 4 findings VideoMAE video 4 findings AlphaFlow 3 findings Amber-7B text 3 findings BakLLaVA text, image 3 findings ChatGLM-6B / ChatGLM-6b-2 text 3 findings Chinchilla 3 findings Claude 2.1 text 3 findings Claude Instant 1 3 findings CodeGen text 3 findings DALL-E text, image 3 findings DALL·E 2 text, image 3 findings DALL·E 3 text, image 3 findings DETR-R50 3 findings DINO-FocalNet-Large 3 findings DeepSeek-2-Chat text 3 findings DeepSeek-2-Coder text 3 findings DeepSeek-V2-0628 text 3 findings ESM3 text 3 findings FoveaBox 3 findings Fuyu text, image 3 findings GCN graph 3 findings GLIDE 3 findings GLM-4V text, image 3 findings GOAT-7B 3 findings GPT-Neo text 3 findings I3D video 3 findings ImageBind-LLM-7B text, image 3 findings LLaMA-Adapter v2 3 findings LLaVA-Med text, image 3 findings LWM-1M-JAX text, video 3 findings LanguageBind text, video 3 findings Lemur-v1-70B / Lemur-70B-Chat-V1 3 findings Llama-VID image, text 3 findings MAE-B/16 image 3 findings MViT V2 image 3 findings Mamba 3 findings MobileNetV3 3 findings Moonshot-v1-8k text 3 findings PaLM 62B 3 findings PandaGPT-7B text 3 findings Phi-2 text 3 findings Phi-3.5 Mini Instruct text 3 findings RegNet image 3 findings RetinaNet-R50 3 findings SlowFast video 3 findings Sybil image 3 findings TD-MPC 3 findings TimesFormer time-series 3 findings Uniformer image 3 findings Video-LLaMA 3 findings VideoCLIP 3 findings X-InstructBLIP-7B text, image 3 findings X3D 3 findings XGen-MM text, image 3 findings AlphaCode 2 findings Aquila-7B 2 findings AutoTikZ / DataTikZ text, image 2 findings BERT-base-cased text 2 findings BEiT image 2 findings BLIP 2 findings Baichuan text 2 findings Baichuan 2 text 2 findings Baichuan2-13B text 2 findings CF2 2 findings CLAP audio, text 2 findings CLIP4Clip 2 findings CLIPBERT 2 findings CONCH image, text 2 findings Cambrian-1 text, image 2 findings Chameleon text, image 2 findings DECAF 2 findings DeepGate3 2 findings DeiT image 2 findings DeiT-III 2 findings DenseNet / DenseNet-101 image 2 findings ESCN 2 findings ESMFlow text 2 findings EdgeNeXt 2 findings EigenFold 2 findings Emu2 text, image 2 findings EquiformerV2 2 findings FLAVA 2 findings Florence-2 image 2 findings GAT graph 2 findings Grounding DINO text, image 2 findings Guanaco 2 findings HRNet 2 findings ImageBind image 2 findings InternLM-2.5-7B text 2 findings InternVL-1.5 text, image 2 findings InternVideo 2 findings LLaVA-Phi text, image 2 findings MLP-Mixer 2 findings Merlot Reserve 2 findings MobileNetV2 2 findings Moondream2 text, image 2 findings Nova Canvas image 2 findings O3 text 2 findings OpenChat-3.5-0106 text 2 findings OpenLLaMA text 2 findings PickScore text, image 2 findings ProGen-2 text 2 findings Qwen-Audio audio, text 2 findings R2D2 text 2 findings RCExplainer 2 findings ReXNet 2 findings STR2STR 2 findings Seed-LLaMA-8B text, image 2 findings StableLM text 2 findings TigerBot-70B-Chat 2 findings TreeRing 2 findings UniPerceiver 2 findings UniVL 2 findings VindLU 2 findings VioLET 2 findings X-CLIP 2 findings XGLM text 2 findings Xlm-R 2 findings Xwin-LM-70B-v0.1 2 findings Zephyr-7B-beta text 2 findings iFlytekSpark-13B text 2 findings mPLUG-2 2 findings 12-in-1 1 finding 3D Gaussian Splatting image 1 finding AASIST audio 1 finding ADM 1 finding AKT 1 finding ALBERT-base-v2 1 finding ALMOST Reward Model 1 finding Agentlm text 1 finding AlexNet image 1 finding AlphaFold2 text 1 finding AnyLoc image 1 finding AutoWebGLM text 1 finding BSRNet 1 finding Bard 1 finding BingChat 1 finding CAIT-S/24 1 finding CLEAR 1 finding CLIPCap 1 finding CLMBR-T-BASE text 1 finding COCA 1 finding CPM-distilled text 1 finding Chat-bison-001 1 finding Chronos time-series 1 finding Claude 1.3 1 finding Claude 3.7 Sonnet text 1 finding CloFNet 1 finding CoMEt 1 finding CoPlace image 1 finding CodeGeex2 1 finding CodeRL 1 finding Codex 1 finding CogAgent text, image 1 finding CogVLM image, text 1 finding ConvIT image 1 finding ConvMixer 1 finding ConvNeXtV2 1 finding ConvNeXtV2-Tiny image 1 finding CycleGAN 1 finding DASR 1 finding DKTF 1 finding DRUM 1 finding DeepFloyd IF text, image 1 finding DeepSeek LLM text 1 finding DeepSeek R1 text 1 finding DeepSeek-VL2 text, image 1 finding DeepSeekMoE text 1 finding Deformable DETR image 1 finding Depth Anything image 1 finding DiT image 1 finding DimeNet++ 1 finding Dreamlike Photoreal 2.0 1 finding E5-V text, image 1 finding EGNN 1 finding ELECTRA-base-discriminator 1 finding ESRGAN 1 finding EVE 1 finding Eurus-RM-7B text 1 finding GEM 1 finding GIN graph 1 finding GKT 1 finding GP-LVM tabular 1 finding GP-UNIT 1 finding GPT-4.1 text 1 finding GPT-4.5 text 1 finding GPT-NeoX-20B text 1 finding GTR-T5-base text 1 finding GVP 1 finding Galactica-6.7B 1 finding GloVe 1 finding GoogLeNet image 1 finding GraphSAGE 1 finding Griffin text 1 finding HLR 1 finding HPS 1 finding HPSv2 1 finding HarmBench Classifier text 1 finding Hawkeye text, video 1 finding HiFaceGAN 1 finding HuggingGPT 1 finding IDDPM image 1 finding ImageReward 1 finding Imagen Video 1 finding Inception image 1 finding InstructPix2Pix 1 finding InstructUIE 1 finding InternLM text 1 finding InternLM2 text 1 finding Internlm2-Reward text 1 finding Jamba text 1 finding LEAPS 1 finding LLaMA-Adapter 1 finding LOVT image, text 1 finding Lambo-2 1 finding LegalBERT text 1 finding Llama-1-7B 1 finding LongVA-7B text, image 1 finding MACE 1 finding MAP-NEO text, image 1 finding MGCA image 1 finding MMRealSR 1 finding MNASNet 1 finding MSA Transformer text 1 finding MagicLens text, image 1 finding Med-Flamingo text, image 1 finding MiDaS image 1 finding Mip-Splatting image 1 finding Mistral Large 2 text 1 finding Mistral Large V2 text 1 finding Mistral-Nemo 12B-Instruct-2407 text 1 finding MoE-LLaVA text, image 1 finding MoGANet 1 finding Moirai time-series 1 finding Molmo text, image 1 finding MolmoE-7B text, image 1 finding Momentor video, text 1 finding NV-Embed-v1 text 1 finding Nova Lite text 1 finding Nova Pro text 1 finding O4-mini text 1 finding OLMoE 6.9B text 1 finding OPUS-MT 1 finding OneLLM text 1 finding OpenAI Moderation text 1 finding OpenJourney 1 finding PGExplainer 1 finding PLIP image, text 1 finding PPOCoder 1 finding PerSAM 1 finding Platypus2-Instruct-70B text 1 finding Prismatic text, image 1 finding ProtoPFormer 1 finding QIKT 1 finding Qwen2-Audio text, audio 1 finding RD-SR 1 finding RS-LDS 1 finding RT-1 1 finding RT-2 1 finding RT-DETR image, text 1 finding RWKV text 1 finding RadFM image 1 finding RealESRNet 1 finding RedPajama text 1 finding RedPajama-INCITE text 1 finding RepVGG image 1 finding ReplkNet-XL 1 finding Reprover 1 finding ResNeXt image 1 finding RexNet 100 1 finding RivaGAN 1 finding SAM 2 image, video 1 finding SAULLM 54B text 1 finding SFR-Embedding-2-R text 1 finding SGC 1 finding SIREN 1 finding SLD-max 1 finding SLD-medium 1 finding SLD-strong 1 finding SLDS 1 finding SLIMG 1 finding SLIP text, image 1 finding SRResNet 1 finding Scaffold-GS image 1 finding SchNet 1 finding SigLIP-2 image, text 1 finding Sketch Transformer 1 finding Skywork text 1 finding Solar 10.7B text 1 finding SpeechGPT text, audio 1 finding SphereNet 1 finding StackLLaMA 1 finding StanfordNLP Reward Model 1 finding Starcoder 1 finding StegaStamp 1 finding Stella-en-1.5b-v5 text 1 finding StyleGAN2-ADA image 1 finding SwinIR 1 finding TAGExplainer 1 finding TAPe 1 finding TimeChat text, video 1 finding TranceptionEVE text 1 finding Tulu 1 finding UForm 1 finding UNITER 1 finding UniIR text, image 1 finding UnifiedQA 1 finding VAST text, image, audio 1 finding VERA text 1 finding VTG-LLM text, video 1 finding Valor text, image, audio 1 finding Van Gogh Diffusion 1 finding ViLA-8B text, image 1 finding ViLBERT 1 finding ViRTex 1 finding ViV1T video, time-series 1 finding VideoChat 1 finding VisualGLM 1 finding WJS text 1 finding WideResNet image 1 finding WizardMath 1 finding XLM-RoBERTa-base 1 finding Xception image 1 finding Xlam-R text, image 1 finding YOLO-World image, text 1 finding YOLOv7 image 1 finding YOLOv8 image, text 1 finding YOLOv9 image 1 finding ZiYA2 text 1 finding code-davinci-002 1 finding mBART-50 1 finding mPLUG-Owl2 text, image 1 finding text-ada-001 1 finding text-ada-002 1 finding ALIGN 0 findings Aegis-Guard-Permissive 0 findings AltCLIP 0 findings Auto-J 0 findings BART-L-MNLI text 0 findings BioBERT 0 findings BioMedCLIP 0 findings CF-GNNExplainer 0 findings Clinical BERT 0 findings CoCondenser text 0 findings Cohere Command 52B 0 findings Command X Large Beta 0 findings Contriever text 0 findings Data Filtering Networks (DfN) 0 findings DeBERTa-V3-NLI text 0 findings Deep Knowledge Tracing 0 findings DiffAb 0 findings DiffAffinity 0 findings DyMEAN graph 0 findings ELLA text, image 0 findings ERNIE-4-8k-0613 0 findings ESM-1F text, graph 0 findings EVA-02 image 0 findings FLUX.1[dev] text, image 0 findings Frozen in Time 0 findings FsfairX-Llama3-RM-v0.1 0 findings GRM-Llama3-8B-SFTReg 0 findings GRU 0 findings Hornet 0 findings Kandinsky v2.2 text, image 0 findings LSTM 0 findings Llama 3.3 70B Instruct 0 findings Megatron-Turing NLG 530B / TNG-L v2 text 0 findings Mistral-Nemo-Instruct text 0 findings MoCo 0 findings ModelScope Text-to-Video 0 findings NATS-Bench 0 findings OpenAssistant/Reward-Model-DeBERTa-V3-Large-V2 text 0 findings PROPEN 0 findings PixArt-α text, image 0 findings Playground v2 text, image 0 findings ProMiM text 0 findings Prometheus 2 text 0 findings PubMedBERT 0 findings QwQ text 0 findings RefineGNN graph, text 0 findings Rotamer Density Estimator / RDE 0 findings STSC-SNN 0 findings Stanford Alpaca 0 findings TA SNN time-series 0 findings TAS-B text 0 findings TinyStories 4-layer 33M / TinyStories 4-layer 33M model text 0 findings UGRNN 0 findings Video Swin Transformer / VideoSwin video 0 findings text-embedding-ada-002 0 findings Findings FX-001 Pairwise Banzhaf interactions explain CLIP similarity more faithfully than single-score methods CLIP / CLIP-ViT (LC) FX-002 FIXLIP gives SigLIP-2 higher pointing-game recognition than CLIP at ViT-B/32 and ViT-B/16 CLIP / CLIP-ViT (LC) , SigLIP , SigLIP-2 FX-003 FIXLIP's strongest interaction in one CLIP example links doll to an image patch reading dollar CLIP / CLIP-ViT (LC) IC-001 Vision-language models perform near chance on the NL-Eye visual abductive reasoning benchmark Gemini 1.5 / Gemini Pro 1.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Claude 3.5 , Claude 3 , LLaVA-NeXT / LLaVA 1.6 , Fuyu , MiniCPM-V , LLaVA-OneVision IC-002 Even when VLMs select the correct hypothesis, their explanations are often invalid or unhelpful Gemini 1.5 / Gemini Pro 1.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Claude 3.5 , Claude 3 , LLaVA-NeXT / LLaVA 1.6 , Fuyu IC-003 HyperDAS dynamically selects intervention tokens and learns linear subspaces in Llama3-8b that mediate entity attributes. Llama 3 IC-004 Retrieval heads are sparse, universal, and causally responsible for long-context retrieval in LLMs Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Yi , Qwen1.5 , Jamba IC-005 GPT-4o underperforms AHA and other VLMs in detecting and reasoning about robotic manipulation failures across multiple datasets. GPT-4o IC-006 The Retriever-Dictionary module improves object detection accuracy of YOLOv7, YOLOv9, Faster R-CNN, and Deformable DETR on COCO 2017 YOLOv7 , YOLOv9 , Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 , Deformable DETR IC-007 Most LLMs do not align closely with human moral preferences on multilingual trolley problems GPT-3 / GPT base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Phi-3 , Qwen 2 IC-008 Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances GPT-4o , Llama 3.1 IC-009 LLM moral preferences show significant language sensitivity but not inequality toward low-resource languages GPT-3 / GPT base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Phi-3 , Qwen 2 IC-010 LLM responses to trolley problems are moderately robust across prompt paraphrases Llama 3 IC-011 Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferences Llama 3.1 , Gemma 2 , Qwen 2 , Llama 2 / Llama 2 base IC-012 Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. Gemma 2 , Llama 3.1 IC-013 Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models. Gemma 2 , Llama 3.1 IC-014 Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses. Gemma 2 IC-015 Truncating MLP weights in Pythia-1b increases the probability of the correct answer in an Indirect Object Identification task Pythia IC-016 Truncating MLP weights in Pythia-1b increases the probability of the correct answer in a factual recall task Pythia IC-017 Truncating MLP weights improves few-shot Chain-of-Thought reasoning accuracy on GSM8K for Phi-3 and Llama-3.1-8B Phi-3 , Llama 3.1 IC-018 Object information is localized to specific visual tokens in LLaVA-1.5 LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Phi IC-019 Visual token representations in LLaVA-1.5 evolve to align with interpretable text tokens LLaVA-1.5 / LLaVA-v1.5 , Qwen2-VL IC-020 LLaVA-1.5 extracts object information directly from visual tokens to the last token in mid-late layers LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Phi IC-021 Vision-language adaptation degrades safety in Llama-2-chat-7b even when training data is filtered for safety Llama 2 / Llama 2 base IC-022 Safety layers in Llama-2-chat-7b show substantial divergence during VL adaptation, correlating with safety degradation Llama 2 / Llama 2 base IC-023 Local scaling, rank, and complexity of Stable Diffusion correlate with generation aesthetics, diversity, and memorization Stable Diffusion IC-024 Reward model trained on local scaling of Stable Diffusion can guide generation to increase diversity and aesthetic scores Stable Diffusion IC-025 LMMs exhibit poor fine-grained perception in locating individual characters on original oracle bones Gemini 1.5 / Gemini Pro 1.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Qwen-VL , XGen-MM , mPLUG-Owl3 , MiniCPM-V , Moondream2 , InternVL2 , GLM-4V , CogVLM2 , LLaVA-NeXT / LLaVA 1.6 , Idefics2 , DeepSeek-VL , InternLM-XComposer2-VL , LLaVA-1.5 / LLaVA-v1.5 IC-026 LLMs can assist in OB rejoining by identifying rejoinable fragments with moderate accuracy, but are not yet truly usable Gemini 1.5 / Gemini Pro 1.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Qwen-VL , GLM-4V , XGen-MM , mPLUG-Owl3 , MiniCPM-V , InternVL2 , LLaVA-NeXT / LLaVA 1.6 , Idefics2 , DeepSeek-VL IC-027 LMM performance in deciphering oracle bone inscriptions is comparable to untrained humans for common characters but declines for rarer and structurally complex characters Gemini 1.5 / Gemini Pro 1.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Qwen-VL , XGen-MM , mPLUG-Owl3 , MiniCPM-V , Moondream2 , InternVL2 , GLM-4V , CogVLM2 , LLaVA-NeXT / LLaVA 1.6 , Idefics2 , DeepSeek-VL , InternLM-XComposer2-VL , LLaVA-1.5 / LLaVA-v1.5 IC-028 SPADE, an abstaining classifier built on top of ResNet, ViT, and VGG models, detects out-of-distribution and adversarial samples with provable guarantees. ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , VGG / VGG13 , ViT IC-029 Large language models show conformity to group answers in multi-agent interactions GPT-3.5 / ChatGPT-3.5 , GPT-4o , Llama 3 , Llama 3.1 , Gemma 2 , Qwen 2 , GLM-4 IC-030 Larger language models exhibit higher independence rates and lower conformity under some protocols Qwen 2 , Llama 3 , Llama 3.1 , Gemma 2 , GPT-4o , GPT-3.5 / ChatGPT-3.5 IC-031 Empowered persona prompts and reflection mechanisms reduce conformity in large language models Llama 3 , Qwen 2 IC-032 Off-policy DPO causes a squeezing effect in LLMs where probability mass shifts to the most confident token, explaining degenerate repetition Pythia , Qwen1.5 IC-033 Pre-training the SFT stage on both chosen and rejected responses mitigates the DPO squeezing effect and improves alignment win rates Qwen1.5 IC-034 Benefit and detriment in RAG can be traded off at token level for Llama-2, OPT and Mistral using representation similarity OPT , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-035 Removing the inductive bias of locality from Vision Transformers improves or matches performance on classification and regression tasks. ViT IC-036 Removing locality from Vision Transformers improves performance in self-supervised learning via Masked Autoencoding. ViT IC-037 Removing locality from Diffusion Transformers improves image generation quality. DiT IC-038 Few embedding dimensions drive the modality gap in CLIP and SigLIP CLIP / CLIP-ViT (LC) , SigLIP IC-039 Object bias in CLIP and SigLIP is not correlated with performance on attribute tasks CLIP / CLIP-ViT (LC) , SigLIP IC-040 Information imbalance triggers both the modality gap and object bias in contrastive VLMs CLIP / CLIP-ViT (LC) IC-041 CLIP and SigLIP use the modality gap to control logit entropy CLIP / CLIP-ViT (LC) IC-042 No single knowledge editing method excels across all criteria when editing visual and user-specific knowledge in LMMs. BLIP-2 , MiniGPT-4 , LLaVA-1.5 / LLaVA-v1.5 IC-043 Five ~7B decoder-only LLMs develop a high-intrinsic-dimensionality phase in intermediate layers that marks the transition from surface-form to abstract linguistic processing, with earlier onset predicting better next-token prediction OPT , Llama 3 , Pythia , OLMo / OLMo base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-044 Tulu-2-13B's internal activations contain a linearly decodable, faithful representation of input-context propositions that persists under prompt injection and backdoor attacks where outputs become unfaithful Tulu 2 , Llama 2 / Llama 2 base IC-045 A 50-dimensional Hessian-identified subspace in Tulu-2-13B causally mediates entity-attribute binding, generalizing to three-entity contexts Tulu 2 , Llama 2 / Llama 2 base IC-046 Tulu-2-13B exhibits gender bias in both its internal binding representation and its outputs, with the output-level bias being stronger than the representation-level bias Tulu 2 , Llama 2 / Llama 2 base IC-047 Tulu-2-13B's entity-attribute binding partially relies on token order as a shortcut, degrading in nested orderings where order and semantic binding conflict Tulu 2 , Llama 2 / Llama 2 base IC-048 Existing video LLMs (TimeChat, VTG-LLM, Momentor, Hawkeye) show limited zero-shot video temporal grounding capability and struggle to improve with fine-tuning TimeChat , VTG-LLM , Momentor , Hawkeye IC-049 GPT-4o shows strong performance on some E.T.Bench event-level tasks (RVQ: 57.7, VHD: 56.9) but very weak performance on others (EPM: 4.5, TAL: 20.0) GPT-4o IC-050 GPT-4o achieves 53.33 overall on Event-Bench with strong event description (57.50) and counter reasoning (63.44) but weaker episodic reasoning (37.33) GPT-4o IC-051 Qwen2-VL (7B) achieves 0.0 on DVC, DVC SLC, and TEM tasks on E.T.Bench Qwen2-VL IC-052 GPT-2 small's IOI circuit activations are linearly decomposable into features for the io, s, and pos attributes, with the l10h0 name mover's attention decomposing into sparse pairwise feature interactions GPT-2 IC-053 In GPT-2 small's l10h0 name mover queries, the io attribute is encoded with higher-magnitude features than the s attribute, and both are causally relevant, but SAEs preferentially learn io features due to the magnitude asymmetry GPT-2 IC-054 Gemma 2 2B performance degrades substantially when routed through Gemma Scope SAEs, and SAE-based feature suppression causes broad cross-domain degradation rather than targeted knowledge removal Gemma 2 IC-055 OLMoE 6.9B experts show no domain specialization, with routing scores evenly distributed across MMLU domains, preventing targeted knowledge unlearning OLMoE 6.9B IC-056 Influence scores of pretraining documents correlate across reasoning queries of the same type, indicating Command R 7B and 35B rely on shared procedural knowledge rather than retrieving specific answers Cohere Command R IC-057 Command R 7B and 35B rely on each individual pretraining document less per nat of generated information for reasoning than for factual questions, with less volatile influence magnitudes Cohere Command R IC-058 The answer to factual questions appears in the top 0.01% most influential pretraining documents for 55% of 7B queries and 30% of 35B queries, but almost never for reasoning questions Cohere Command R IC-059 Code data is strongly overrepresented in the most influential pretraining documents for reasoning queries in Command R 7B and 35B Cohere Command R IC-060 SAE features in Pythia-160m and Mamba-130m exhibit high cross-architecture similarity with a depth-scaled correspondence Pythia , Mamba IC-061 The induction circuit in Mamba-130m is structurally analogous to the transformer induction circuit, with an off-by-one motif in SSM state writing Mamba IC-062 Llama 3 70B implements temporal difference learning in-context for reward-based RL, with causally relevant SAE features in its residual stream, while Llama 3 8B performs at chance Llama 3 IC-063 Llama 3 70B learns global graph structure via TD learning, building successor-representation-like geometry in its residual stream that is causally supported by TD latents Llama 3 IC-064 The TD learning mechanism identified in Llama 3 70B generalizes to Gemma-2-27B and Qwen-2.5-72B across all three tasks Gemma 2 , Qwen2.5 IC-069 EVA-CLIP's dense patch features are semantically contaminated by surrounding context, degrading their spatial quality EVA-CLIP IC-070 Region-language alignment fine-tuning degrades EVA-CLIP's spatial awareness as measured by unsupervised segmentation EVA-CLIP IC-071 DINOv2's dense features are dominated by global context, impairing fine-grained spatial detail DINOv2 IC-072 LLM agents of varying scales exhibit a failure mode on web automation tasks when processing raw, complex web page observations, with the penalty being more severe for smaller models GPT-4o , Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 , Llama 3.1 IC-073 Released LLMs (GPT-4o, Llama-3.1-70B, Qwen2-7B, etc.) show limited workflow orchestration capability that degrades as workflow complexity increases GPT-4o , Llama 3.1 , Qwen 2 IC-074 Released LLMs achieve F1 plan scores between 42.7 and 86.7 on the T-Eval plan task GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Qwen , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3.1 , Llama 2 / Llama 2 base , Vicuna , Baichuan2-13B , Wizardlm , Qwen1.5 IC-075 GPT-4o-mini and Qwen2.5-72B achieve low precision and recall when used as API retrievers for workflow orchestration GPT-4o , Qwen2.5 IC-076 GoogLeNet and ViT exhibit input space mode connectivity: inputs with similar predictions are connected by low-loss paths, with real-real pairs showing approximately linear paths and real-adversarial pairs showing significantly higher barriers GoogLeNet , ViT IC-077 VGG-16's loss landscape barrier height distinguishes adversarial from real inputs, enabling a detection method that outperforms baselines on DeepFool and C&W attacks VGG / VGG13 IC-078 GPT-4's self-verification loop causes performance collapse due to high false negative rates in binary verification GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-079 GPT-4's free-form critique generation is unreliable, containing hallucinated edges, vertex colors, and precondition states GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-080 GPT-4's performance is largely insensitive to the content of feedback; simple re-prompting with a sound verifier (sampling) matches or exceeds detailed critique GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-081 ViT/L-16 exhibits lower sensitivity to token-wise Gaussian perturbations than ConvNeXtV2-Tiny on ImageNet-1k ViT , ConvNeXtV2-Tiny IC-082 GPT-3.5, GPT-4o, Claude-3.5-Sonnet, and Llama-3.1-8B produce explanations on the BBQ social bias task that are systematically unfaithful for identity and behavior concepts while remaining faithful for context concepts, with specific patterns of hiding safety-measure influence and social bias GPT-3.5 / ChatGPT-3.5 , GPT-4o , Claude 3.5 , Llama 3.1 IC-083 GPT-3.5, GPT-4o, and Claude-3.5-Sonnet produce unfaithful explanations on MedQA medical questions, omitting high-effect clinical concepts such as the patient's mental status while over-referencing low-effect concepts GPT-3.5 / ChatGPT-3.5 , GPT-4o , Claude 3.5 IC-084 Safety alignment in Llama-2-7b-chat and Gemma-7b-1.1-it is shallow, with the KL divergence from the base model concentrated in the first few output tokens, making the models vulnerable to prefilling attacks Llama 2 / Llama 2 base , Gemma IC-085 Unaligned base models Llama-2-7b and Gemma-7b produce predominantly safe continuations when prefilled with refusal prefixes, demonstrating a pre-existing safety shortcut Llama 2 / Llama 2 base , Gemma IC-086 Fine-tuning Llama-2-7b-chat on 100 harmful examples for 6 gradient steps increases the attack success rate from 1.5% to 87.9%, with per-token dynamics showing the distributional change concentrated in the first few tokens Llama 2 / Llama 2 base IC-087 Answer symbol production in OLMo 7B Instruct, Llama 3.1 8B Instruct, and Qwen 2.5 1.5B Instruct is causally attributed to a few middle layers and specifically their multi-head self-attention mechanisms, with a sparse set of 1-4 attention heads per layer responsible OLMo / OLMo base , Llama 3.1 , Qwen2.5 IC-088 OLMo 7B Instruct and Qwen 2.5 1.5B Instruct exhibit a two-stage process for unusual answer symbols, initially assigning non-negligible probability to expected symbols (a/b/c/d) before switching to the actual prompt symbols at a specific later layer OLMo / OLMo base , Qwen2.5 IC-089 OLMo 0724 7B base learns formatted multiple-choice question answering between 80k and 100k training steps, transitioning from near-random to near-perfect accuracy on the synthetic colors task OLMo / OLMo base IC-092 All evaluated LLMs show consistent F1 degradation to at most 0.60 when two or more events match a retrieval cue GPT-4o , Claude 3 , Claude 3.5 , Llama 3.1 , O1 / OpenAI-o1-preview IC-093 No evaluated LLM achieves perfect confabulation avoidance on questions about non-existent events GPT-4o , Claude 3 , Claude 3.5 , Llama 3.1 , O1 / OpenAI-o1-preview IC-094 Episodic recall accuracy degrades systematically from content cues to space cues to time cues across all evaluated LLMs GPT-4o , Claude 3 , Claude 3.5 , Llama 3.1 , O1 / OpenAI-o1-preview IC-095 Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological ordering GPT-4o , Claude 3 , Claude 3.5 , Llama 3.1 , O1 / OpenAI-o1-preview IC-096 Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the top Llama 3.1 , Qwen 2.5 72B Instruct , O3 , GPT-4o , Llama 3 , GPT-4.5 , DeepSeek R1 , Claude 3.7 Sonnet IC-097 Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3B Llama-3.2-3B , Llama 3.1 IC-098 LegalBERT performs below the constant classifier baseline on CASELAWQA due to its 512-token context window LegalBERT , SAULLM 54B IC-099 GPT-4 and Claude 3 Opus can be prompted to selectively underperform on WMDP while maintaining general performance on MMLU and CSQA GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 IC-100 GPT-4, GPT-3.5, Claude 3, Llama 3 8B, and Llama 3 70B can be prompted to approximately calibrate their accuracy to specific target percentages on MMLU GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Claude 3 , Llama 3 IC-1003 Pretrained ResNet-50 and ViT-B/16 exhibit neuron activation patterns that are separable between in-distribution and out-of-distribution inputs, enabling post-hoc OOD detection without model modification ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , ViT IC-1005 Social bias neurons in BERT-base-cased and RoBERTa-base are concentrated in the deepest transformer layers BERT , RoBERTa / RoBERTa-L IC-1007 LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback ChatGPT , GPT-3 / GPT base , Llama 2 / Llama 2 base IC-1008 The magnitude of CRITIC's improvement on mathematical program synthesis scales with Llama-2 model size Llama 2 / Llama 2 base IC-101 GPT-4 and Claude 3 struggle to emulate a lower capability profile (high school freshman level) via zero-shot prompting, with only moderate improvement from chain-of-thought prompting GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 IC-1011 OpenAI CLIP loses approximately 8% zero-shot retrieval accuracy on 2021–2022 data compared to OpenCLIP models trained on data through 2022, while standard benchmarks show no such gap CLIP / CLIP-ViT (LC) , OpenCLIP IC-1012 Connected regions in the latent space of Stable Diffusion v2.1, v1.5, and GLIDE produce distorted images independent of the text prompt Stable Diffusion , GLIDE IC-1013 Latent samples in Stable Diffusion v2.1, v1.5, and GLIDE can produce images of associated backgrounds rather than the key object, with failure rates of 2.7%, 9.2%, and 50.5% under random sampling respectively Stable Diffusion , GLIDE IC-1014 A single adversarial token embedding appended to any input prompt overwrites the prompt in Stable Diffusion v2.1 to generate a target object, with CLIP similarity to the original prompt (0.742) remaining higher than to the target (0.546) Stable Diffusion , GLIDE IC-1015 GPT-J and 10 other LLMs exhibit overthinking: calibrated accuracy given incorrect few-shot demonstrations peaks at a critical layer then declines, and ablating 5 false induction heads in late layers reduces the accuracy gap by 38.9% on average GPT-J , GPT-2 , GPT-NeoX-20B , Pythia , Llama 2 / Llama 2 base IC-1016 ImageNet-pretrained ResNet-50 backbone exhibits shortcut bias toward background features when adapted to bird classification via a new readout layer ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-1017 Vision and language models pre-trained on noisy data exhibit degraded OOD transfer that is partially recoverable via SVD-based feature-space regularization EfficientNet , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , Swin Transformer , ViT , ConvNeXt , BERT , RoBERTa / RoBERTa-L , GPT-2 , text-ada-002 IC-102 GPT-4o's spatial understanding degrades when depth maps are provided as additional input on SpatialBench GPT-4o IC-1023 The binary activation pattern of standard CNNs (VGG, ResNet) carries most of the classification information, as shown by APoP VGG / VGG13 , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-103 LLMs' value rankings align with the universal human value hierarchy under most prompting conditions GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.0 Pro , Llama 3.1 , Gemma 2 IC-1035 Reprover achieves 0% accuracy on sorry theorems in advanced mathematics repositories (PFR, Hairy Ball Theorem, Coxeter) while proving basic theorems in other repositories Reprover IC-104 Value anchor prompting produces LLM value correlation structures that closely match the human circular value structure, while standard prompting does not GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.0 Pro , Llama 3.1 , Gemma 2 IC-1044 Counterfactual GNN explainers produce statistically infeasible recourses that violate topological constraints in molecular datasets RCExplainer , CF2 , CLEAR IC-1045 Factual GNN explanations do not capture the full data signal: retraining on explanations fails to reproduce predictions while retraining on residuals preserves them PGExplainer , TAGExplainer , GEM , CF2 , RCExplainer IC-105 Value anchoring produces a sinusoidal scoring pattern around the circular value structure, with scores decreasing as circular distance from the anchor increases GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.0 Pro , Llama 3.1 , Gemma 2 IC-1050 Released models generate patches that are less than half the length of gold solutions and rarely edit more than one file GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1053 Instruction tuning suppresses in-context learning in LLaMA, Vicuna, and OPT-IML, with the suppression being largest for English prompts and partially recoverable via translation to other languages LLaMA , Alpaca , Vicuna , OPT IC-1054 Code fine-tuning degrades English natural language reasoning in Code LLaMA relative to LLaMA-2, but the effect is negligible or slightly positive in French, Spanish, and German Llama 2 / Llama 2 base , Code Llama IC-1055 Safety fine-tuning suppresses harmful content generation in ChatGPT relative to GPT-3.5, but the suppression is substantially weaker for non-English prompts GPT-3.5 / ChatGPT-3.5 , ChatGPT IC-106 Logit lens on LLaVA and InstructBLIP image representations shows higher internal confidence for objects present in the image than for hallucinated objects LLaVA-1.5 / LLaVA-v1.5 , InstructBLIP , LLaVA-NeXT / LLaVA 1.6 , Cambrian-1 IC-1069 Stable Diffusion v2.1's compositional understanding on the ARO benchmark is significantly higher than previously reported by MMSE-based scoring Stable Diffusion , OpenCLIP IC-107 Linear orthogonalization of LLaVA and InstructBLIP image features against text embeddings removes hallucinated objects at 83-86% individual rate versus 7-16% for correctly detected objects LLaVA-1.5 / LLaVA-v1.5 , InstructBLIP , LLaVA-NeXT / LLaVA 1.6 , Cambrian-1 IC-1070 In Stable Diffusion v2.1, attention maps do not reliably predict the effect of prompt interventions on generated images, while conditional mutual information does Stable Diffusion IC-1071 Stable Diffusion v2.1's pixel-wise conditional mutual information localizes abstract words (adjectives, adverbs, verbs) more effectively than attention, but is less effective than attention for object segmentation Stable Diffusion IC-1078 LLaMA models (7B through 65B) exhibit gender bias in language generation, coreference resolution, and sentence likelihood, with stereotypical associations driving predictions LLaMA IC-1079 Causal tracing reveals that mid-upper MLP layers (layers 18–25 in 7B) are the primary mediators of stereotypical gender bias in LLaMA, while the last layers show negative coefficients that counter the bias LLaMA IC-108 Per-patch logit lens confidence in LLaVA localizes objects spatially, achieving mAP 79.90 on ImageNet segmentation, 8.03% above raw VLM attention LLaVA-1.5 / LLaVA-v1.5 IC-1080 OPUS-MT small, OPUS-MT large, and mBART-50 in their default (unfine-tuned) form achieve low context-sensitive disambiguation accuracy on discourse-level phenomena OPUS-MT , mBART-50 IC-1086 RS LDS fails to increase its number of active states when the underlying dynamics change non-stationarily RS-LDS IC-1087 SLDS produces poor dynamical accuracy on the NASCAR task because it lacks recurrent switching SLDS IC-1088 ICL predictions in LLaMA, LLaMA-2, and Falcon models depend on in-context label information and can learn truly novel label relationships Llama 2 / Llama 2 base , LLaMA , Falcon IC-1089 ICL in LLaMA, LLaMA-2, and Falcon models cannot fully overcome pre-training label preferences when in-context labels are flipped Llama 2 / Llama 2 base , LLaMA , Falcon IC-109 GPT-3.5-turbo and GPT-4 produce cycles in inferred causal graphs when using pairwise prompts, with cycle counts growing sharply on larger graphs GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1090 ICL in LLaMA, LLaMA-2, and Falcon models preferentially uses in-context label information closer to the query rather than treating all examples equally Llama 2 / Llama 2 base , LLaMA , Falcon IC-1099 AdamW-pretrained vision models (ViTs, ConvNeXt) have disproportionately large embedding-layer gradients at initialization, causing SGD fine-tuning to degrade OOD accuracy by up to 15% relative to AdamW CLIP / CLIP-ViT (LC) , ViT , DINO , ConvNeXt , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-110 Phi-3 (3.8B) and Llama-3 (8B) with triplet prompting outperform GPT-4 with pairwise prompting on causal graph orientation Phi-3 , Llama 3 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1105 PAC-Bayes generalization bounds for discrete class prompts on CLIP are within a few percentage points of the actual test error across CIFAR-10, CIFAR-100, ImageNet, FMOW, and OfficeHome CLIP / CLIP-ViT (LC) IC-1106 CLIP prompts found by greedy search do not fit random labels: train and test error drop in tandem as the fraction of flipped labels increases, unlike a linear probe which achieves near-random accuracy CLIP / CLIP-ViT (LC) IC-111 ViT-B/16 pretrained with MAE exhibits higher attention diversity than ViT-B/16 pretrained with MoCo v3, DINO, or DeiT ViT IC-112 Released LLMs show a reproducible failure mode where numerical task accuracy degrades sharply as input digit length increases GPT-4o , Llama 3.1 , Llama 2 / Llama 2 base , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Qwen 2 IC-1123 All seven published concept erasure methods applied to Stable Diffusion 1.4 can be circumvented by learned word embeddings, demonstrating that targeted concepts are input-filtered rather than truly removed from the model Stable Diffusion IC-1127 The LM head in GPT-2, GPT-J, BLOOM, Pythia, and LLaMA-2 projects all input token hidden states into interpretable token distributions over the vocabulary, and these distributions converge approximately monotonically toward the final layer's distribution GPT-2 , GPT-J , BLOOM , Pythia , Llama 2 / Llama 2 base IC-113 Released LLMs show a reproducible failure mode where accuracy on fraction and scientific notation tasks falls below 20% even for the shortest inputs GPT-4o , Llama 3.1 , Llama 2 / Llama 2 base , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Qwen 2 IC-1133 LLMs are highly receptive to coherent counter-memory as sole evidence, contradicting prior findings of stubbornness with entity-substitution counter-memory ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , PaLM 2 , Qwen , Llama 2 / Llama 2 base , Vicuna IC-1134 LLMs show strong confirmation bias in multi-source settings, preferring evidence consistent with parametric memory, with stronger bias for popular entities ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , PaLM 2 , Qwen , Llama 2 / Llama 2 base , Vicuna IC-1135 LLMs show order sensitivity to evidence position in context, with PaLM2 and LLaMA2-7B showing memorization ratio variations exceeding 30% ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , PaLM 2 , Llama 2 / Llama 2 base IC-1136 Larger LLMs (LLaMA2-70B, Vicuna-33B) are more stubborn than their smaller counterparts (LLaMA2-7B, Vicuna-7B) when encountering incoherent entity-substitution counter-memory Llama 2 / Llama 2 base , Vicuna IC-1137 Stable Diffusion represents certain concepts primarily through specific named exemplars rather than abstract category features Stable Diffusion IC-1138 Stable Diffusion simultaneously encodes multiple meanings of homograph concepts in a single representation Stable Diffusion IC-1139 Stable Diffusion encodes social biases in its internal concept representations that are not always visually apparent Stable Diffusion IC-114 Released LLMs cannot reliably identify a specific digit in a number as the number's length increases, with GPT-4o achieving only 20% on get-digit in the xl range GPT-4o , Llama 3.1 , Llama 2 / Llama 2 base , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Qwen 2 IC-1140 Stable Diffusion's internal concept representations encode visual and structural similarities (shape, texture, color) that transcend textual semantics Stable Diffusion IC-1144 GPT-4 and GPT-3.5 produce high rates of irrelevant (fabricated) books when answering constraint queries from parametric knowledge, with a sharp phase transition at low author popularity GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1145 Providing complete context eliminates irrelevance but does not fix constraint satisfaction for GPT-4 or GPT-3.5 GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1146 Self-context (self-retrieval) chain-of-thought increases the rate of fabricated books compared to no-context for both GPT-4 and GPT-3.5 GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1147 GPT-4 outperforms GPT-3.5 on all KITAB metrics but the gap is modest, with all-correctness below 35% for both, suggesting scale alone does not resolve constraint satisfaction GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1148 ChatGPT and InstructGPT exhibit positional bias in their reliance on in-context examples, with ChatGPT showing decreasing attention by position and InstructGPT showing a U-shaped pattern ChatGPT , InstructGPT IC-1149 ChatGPT and Llama-2-7b-chat fail to recognize unanswerable questions on SQuAD 2.0, with Llama-2-7b-chat scoring only 3.72% accuracy on no-answer questions ChatGPT , Llama 2 / Llama 2 base IC-115 NUPA performance is largely independent of model size within a family: GPT-4o and GPT-4o-mini show nearly identical performance, as do Qwen2-72B and Qwen2-7B GPT-4o , Qwen 2 IC-1150 ChatGPT and Llama-2-7b-chat underperform humans by 20 and 31 points respectively on out-of-distribution NLU tasks in GLUE-X ChatGPT , Llama 2 / Llama 2 base IC-1151 LLaMA, OPT, LLaMA-2, Mistral, and GPT-J all exhibit token co-occurrence reinforcement, where the probability of generating a token increases monotonically with the number of its contextual co-occurrences LLaMA , OPT , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , GPT-J IC-1152 Token reinforcement in demonstrations constrains LLaMA-65B's output to valid label spaces on MMLU and enables chain-of-thought pattern following on GSM8K without requiring question content LLaMA IC-1153 Intentionally constructed spurious token connections in MMLU demonstrations misdirect LLaMA-65B's in-context learning toward specific answer choices LLaMA IC-1154 LLaMA-65B exhibits a selection bias where zero-shot accuracy varies substantially by answer choice, with 'a' at 71.58% and 'd' at 52.28% LLaMA IC-1155 ViT-B/16 (ImageNet-21k) fine-tuned with VPT outperforms full fine-tuning on 16 of 19 VTAB-1k tasks, with the advantage concentrated in high-task-disparity and similar-distribution scenarios and narrowing as downstream data grows ViT , Swin Transformer IC-1156 The VPT advantage over FT for ViT-B/16 is not explained by overfitting resistance or additional optimization dimensions; the specific feature-preservation mechanism of VPT is the key factor ViT IC-1157 GPT-4 and other LMs show a large gap between rule induction and rule application, with task accuracy dropping to near zero on MiniScan when the LM itself applies its own proposed rules GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1158 GPT-4 and other LMs are brittle to noisy exemplars and unfamiliar output representations, with performance degrading sharply even under minimal perturbation GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1159 GPT-4 is a strong inductive hypothesis proposer, achieving high accuracy on inductive reasoning benchmarks when its generated rules are applied by a symbolic interpreter GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base IC-116 In LLaVA-7B, multi-head attention modules drive hallucination more than MLP modules, and targeted intervention on specific hallucination heads reduces the hallucination rate by up to 1.7x LLaVA-1.5 / LLaVA-v1.5 , MiniGPT-4 IC-1165 ProtoPFormer achieves 42.2% attribute identification accuracy on CUB-200-2011 in a 7-rater human evaluation ProtoPFormer IC-1166 MPLUG-Owl's VQA accuracy on VQA-X increases from 68.30% to 74.48% when prompted with progressively higher-quality rationales generated by RAPPER mPLUG-Owl IC-1167 GPT-4 Code Interpreter's mathematical reasoning accuracy is positively correlated with code usage frequency, with its iterative code generation and self-debugging mechanism as the primary driver of its 69.69% zero-shot MATH accuracy GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1168 For GPT-4 Code Interpreter, code-based self-verification improves MATH accuracy to 73.54% while natural language self-verification slightly degrades it to 69.29% relative to the 69.69% base GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1169 CodeLlama-7B and CodeLlama-34B show improvement with CSV prompting on GSM8K and MATH but at much lower absolute accuracy than GPT-4 Code Interpreter CodeLlama-13B , CodeLlama-34B IC-117 Hallucination heads in LLaVA-7B and MiniGPT-4 are concentrated in the middle and deeper layers of the transformer LLaVA-1.5 / LLaVA-v1.5 , MiniGPT-4 IC-1170 GPT-3.5, Llama2, PaLM2, and GPT-4 are susceptible to a CoT-prompting backdoor attack (BadChain) on complex reasoning tasks, with stronger reasoning models showing higher attack success rates GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , PaLM 2 , Llama 2 / Llama 2 base IC-1171 Pre-trained ViT, MAE, and ResNet50 (supervised and MoCo v2) place visually similar but semantically distinct ImageNet classes (mop, broom, puck, crutch) in close proximity in their feature space ViT , MAE , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-1175 ChatGPT can be prompted to generate misinformation with near-perfect success for implicit methods but is largely resistant to explicit misinformation requests ChatGPT IC-1176 ChatGPT-generated misinformation is harder for humans to detect than human-written misinformation with the same semantics ChatGPT IC-1177 LLM-generated misinformation is harder for LLM detectors to detect than human-written misinformation with the same semantics ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 2 / Llama 2 base , Vicuna IC-118 Hallucination heads in LLaVA-7B and MiniGPT-4 allocate 4.75x more attention to text tokens than image tokens, and this pattern is inherited from the base language model LLaVA-1.5 / LLaVA-v1.5 , Vicuna , MiniGPT-4 , Llama 2 / Llama 2 base IC-1188 GPT-3 procedural planning performance is scale-dependent, with Curie (6.7B) scoring 3.75 and Davinci (175B) scoring 4.90 overall quality in few-shot settings, while GPT-4 achieves 4.81 overall and 5.00 order GPT-3 / GPT base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-119 The number of salient hallucination heads decreases as model size increases within the LLaVA family LLaVA-1.5 / LLaVA-v1.5 , LLaVA-NeXT / LLaVA 1.6 IC-1193 Vicuna and Alpaca achieve 0% pass rate on all ToolBench tool-use instructions, while GPT-4 and ChatGPT reach 71.1% and 64.8% with DFSDT, revealing a wide capability gap in tool use among released LLMs Vicuna , Alpaca , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , ChatGPT , GPT-3 / GPT base IC-1194 GPT-3.5-turbo, GPT-4, and GPT-3.5-turbo-0613 exhibit 50-58% inconsistency between their ratings and rankings feedback on the same response pairs GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1195 Model substitution adversarial attack reduces TreeRing AUROC to 0.14 at ε=2/255 and StegaStamp AUROC to 0.492 at ε=12/255 TreeRing , StegaStamp IC-1196 Blending a watermarked noise image with a clean image causes watermark detectors to falsely flag clean images as watermarked RivaGAN , TreeRing IC-1199 GPT-2 next-token distributions contain correctable tail errors from the softmax bottleneck that degrade generation quality under low-entropy sampling, with basis-aware threshold sampling improving MAUVE across all four sizes GPT-2 IC-120 YOLO-World can replace SAM as a 2D crop generator in OpenMask3D's pipeline with nearly equivalent mAP but 1.76x faster inference YOLO-World , SAM , YOLOv8 , RT-DETR IC-1200 GPT-2-XL's untruncated next-token log-probability matrix has rank saturating at its hidden dimensionality of 1600, while truncation sampling produces post-truncation distributions whose estimated rank grows far beyond 1600 GPT-2 IC-1201 CLIP-ViT-L/14 image features support 200-way zero-shot EEG-based object recognition better than ViT-B/16 or ResNet-50 features when used as a frozen encoder in a contrastive learning framework CLIP / CLIP-ViT (LC) , ViT , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , EVA-CLIP IC-1202 Stable Diffusion 1.5 and 2.1 exhibit a cropping failure mode where synthesized objects are cut off at image boundaries Stable Diffusion IC-1203 Stable Diffusion 1.5 and 2.1 receive low user-preference win rates (7.91% and 6.71%) in a four-way comparison against SDXL Stable Diffusion IC-1204 SD-VAE 1.x and SD-VAE 2.x achieve lower reconstruction quality than the new SDXL-VAE on COCO 2017 Stable Diffusion IC-1205 GPT-3 models (ada, curie, davinci) achieve near-zero accuracy on zero-shot arithmetic tasks but learn them rapidly with 1000 fine-tuning samples GPT-3 / GPT base IC-1206 GPT-2-XL, GPT-J, Falcon-7B, Llama-2-7B, and Llama-2-13B are vulnerable to backdoor injection via lightweight parameter editing with only 15 samples, achieving near-100% attack success rate while preserving clean performance GPT-2 , GPT-J , Falcon , Llama 2 / Llama 2 base IC-1207 For GPT-2-XL, backdoor injection via parameter editing is most effective on intermediate layers (15-35) and notably less effective on the first 10 and last 5 layers GPT-2 IC-1208 Larger LLMs (Llama-2-13B) require more data samples for successful backdoor injection via parameter editing compared to smaller models (GPT-2-XL 1.5B) GPT-2 , Llama 2 / Llama 2 base IC-121 Gemini 1.0 Pro, when prompted as a zero-shot chain-of-thought judge with majority voting, underperforms fine-tuned smaller Gemma models as verifiers on GSM8K Gemini 1.0 Pro IC-1217 LLaMA 65B's token-probability readout fails to capture human decision-making, producing near-chance NLL and no human-like exploration behavior LLaMA IC-1218 GPT-4 achieves 59.72% accuracy on choices13k and 80.3% on the horizon task when modeling human decisions GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1219 Frozen GPT-2 XL contains pre-trained attention heads that implement the nearest-neighbor algorithm GPT-2 IC-122 Concept representations in Llama-2-7B, Gemma-7B, and Llama-2-13B become more consistent in deeper layers Llama 2 / Llama 2 base , Gemma IC-1220 GPT-4, GPT-3.5-turbo, and Llama-2-70B can implement learning algorithms in-context on novel boolean functions, competing with nearest-neighbor baselines GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base IC-1221 LLM performance on in-context boolean function learning is scale-dependent, with GPT-2 failing and Llama-2 models improving gradually with size GPT-2 , Llama 2 / Llama 2 base IC-1225 GPT-4 generates realistic dynamic scene layouts from text prompts with only 3 in-context examples, achieving 98% average accuracy across 5 spatiotemporal tasks, with physics knowledge (gravity, elasticity, perspective) generalising to unseen objects from its weights GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1226 SAM's ViT-B encoder achieves 54.2% ImageNet-1k linear probing accuracy versus 67.7% for MAE's ViT-B, indicating its segmentation pretraining impairs high-level semantic representation SAM , MAE IC-1227 SAM's segmentation pretraining shifts attention heads toward local focus in deeper layers, unlike its MAE initialization which retains global attention throughout SAM , MAE IC-123 Llama-2-7B, Gemma-7B, and Llama-2-13B organize 16 concepts into hierarchical clusters in their representation space that reflect real-world category structure Llama 2 / Llama 2 base , Gemma IC-1231 On APPS, released code generation models span pass@1 from 0.20 (GPT-3 175B) to 6.20 (CodeRL), with value-based and policy-based RL methods outperforming supervised baselines CodeRL , GPT-2 , GPT-3 / GPT base , GPT-Neo , GPT-J , Codex , AlphaCode , PPOCoder IC-1232 GPT-2 XL and GPT-J exhibit knowledge conflict when subjected to reverse and composite knowledge edits, with ROME and MEMIT showing near-total failure on reverse edits GPT-2 , GPT-J IC-1233 GPT-2 XL and GPT-J exhibit irreversible knowledge distortion after round-editing, with the effect being more severe when the edit target is semantically distant from the true labels GPT-2 , GPT-J IC-1236 Among 7B LLMs, Llama-2-7b achieves the best zero-shot COMET scores in both translation directions, while MPT-7b leads in BLEU for en-to-xx Llama 2 / Llama 2 base , MPT , OPT , BLOOM , Falcon , Llama-1-7B , GPT-3.5 / ChatGPT-3.5 IC-1237 Llama-2-7b's pre-existing translation knowledge is diluted by large amounts of parallel data, causing COMET to decline after 100k examples Llama 2 / Llama 2 base , MPT IC-1238 Llama-2-13b produces off-target non-translation outputs in zero-shot English-to-foreign-language translation Llama 2 / Llama 2 base IC-1239 Llama-2-7b and Llama-2-13b achieve top zero-shot cross-lingual performance among 7B models on XNLI, XStoryCloze, and XWinograd Llama 2 / Llama 2 base , Xlm-R , XGLM , BLOOM , MPT IC-124 Direct comparison of CLIP image embeddings with CLAP audio embeddings achieves near-chance retrieval, while logsumexp bridging through the shared language modality recovers 62% recall@10 on AudioSet CLIP / CLIP-ViT (LC) , CLAP IC-1240 GPT-4-0314 achieves 85% zero-shot accuracy on situational-awareness questions about its own architecture and training GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1241 GPT-4 (14 March 2023) achieves 100% zero-shot accuracy at classifying whether news articles could be part of its pre-training data GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1242 GPT-4 wrote a working script that called an instance of itself on its API as part of a plan to gain internet access GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1243 GPT-4 and GPT-4 Turbo achieve near-zero scores on GAIA level 3 and single-digit to low-double-digit scores on levels 1-2, compared to 87-94% for human annotators GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1244 GPT-4's non-zero scores on GAIA web browsing questions are largely due to memorization of intermediate information from training data rather than actual web browsing GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1249 TD-MPC exhibits training instability and performance degradation when the planning horizon is set to 20 time steps TD-MPC IC-125 LanguageBind's direct evaluation achieves 70% recall@10 on AudioSet, empirically validating that the inner product between unpaired modality representations recovers the correct probability ratio LanguageBind IC-1250 GPT-2-medium shares 78% of its top attention heads between the IOI circuit and the colored objects circuit GPT-2 IC-1251 Intervening on four attention heads in GPT-2-medium boosts colored objects accuracy from 49.6% to 93.7% by making the circuit behave like the IOI circuit GPT-2 IC-1252 Circuit overlap between IOI and colored objects in GPT-2 decreases as model scale increases from medium to xl GPT-2 IC-1256 MPT-7B-Chat produces non-committal responses rather than proper refusals on unsafe instructions MPT IC-1257 Guanaco acknowledges the illegality of requested actions but still provides the harmful information Guanaco IC-126 CLIP and CLAP language representations are statistically indistinguishable from a uniform distribution on the hypersphere CLIP / CLIP-ViT (LC) , CLAP IC-1261 LLaMA-2 attention to constraint tokens correlates with factual correctness, and a linear probe on these attention weights predicts factual errors comparably to model confidence Llama 2 / Llama 2 base IC-1262 LLaMA-2 factual query accuracy improves with entity popularity and decreases with query constrainedness, with larger models showing better performance on less popular and more constrained queries Llama 2 / Llama 2 base IC-1263 LLaMA-2 7B and 13B attention signal for predicting factual errors is available by approximately 50% of layers, enabling early stopping without performance degradation, while LLaMA-2 70B shows a slight performance drop Llama 2 / Llama 2 base IC-1264 LLMs are overconfident when verbalizing confidence, with values concentrated in 80–100% and multiples of 5, yielding high ECE across all five tested models GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Vicuna , Llama 2 / Llama 2 base IC-1265 Calibration and failure prediction improve as model capability scales from GPT-3 to GPT-4, but remain far from ideal GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Vicuna , Llama 2 / Llama 2 base IC-1266 For GPT-3, white-box token-probability methods outperform black-box verbalized confidence in uncertainty estimation, but the gap is narrow (0.522–0.605 AUROC) and both remain near random GPT-3 / GPT base IC-1267 LLMs' alignment with human privacy judgments drops sharply as contextual complexity increases from tier 1 to tier 3 GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , ChatGPT , InstructGPT , Mixtral , Llama 2 / Llama 2 base IC-1268 LLMs leak private information in theory-of-mind scenarios even when explicitly instructed to preserve privacy GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , ChatGPT , InstructGPT , Mixtral , Llama 2 / Llama 2 base IC-1269 LLMs leak secrets to inappropriate recipients in meeting summarization and action-item generation tasks GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , ChatGPT , InstructGPT , Mixtral , Llama 2 / Llama 2 base IC-127 ImageBind's direct evaluation closely matches logsumexp for both vision-language and audio-language alignment, validating the law for ImageBind ImageBind IC-1270 Chain-of-thought prompting does not mitigate privacy leakage in GPT-4 or ChatGPT GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , ChatGPT IC-128 Released VLMs exhibit sycophancy, agreeing with incorrect user opinions while ignoring visual evidence, with LLaVA-1.5 showing the highest rate (94.6%) and InternLM-XComposer2-VL-1.8B the lowest (28.8%) BLIP-2 , InstructBLIP , LLaVA-1.5 / LLaVA-v1.5 , mPLUG-Owl2 , InternVL-1.5 , InternLM-XComposer2-VL , Gemini , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1283 CLIP's InfoNCE training objective is mathematically equivalent to performing generalized spectral clustering on the bipartite image-text pair graph CLIP / CLIP-ViT (LC) IC-1284 Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on as few as 10 explicitly harmful examples removes their safety alignment, raising harmfulness rates to 80-92% GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base IC-1285 Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on 10 implicitly harmful identity-shifting examples (containing no toxic content) jailbreaks their safety alignment GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base IC-1286 Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on benign utility-oriented datasets (Alpaca, Dolly, LLaVA-Instruct) degrades their safety alignment without any malicious intent GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base IC-1287 A backdoor can be implanted in GPT-3.5 Turbo via fine-tuning that is undetectable by standard safety auditing: the model appears safe on plain prompts but fulfills harmful instructions when a 3-word trigger is appended GPT-3.5 / ChatGPT-3.5 IC-1288 GPT-4, ChatGPT, and GPT-4V fail to close the human-machine gap on Bongard-OpenWorld, with InstructBLIP captions differentially degrading ChatGPT while improving GPT-4 ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , BLIP-2 , InstructBLIP IC-1289 OpenFlamingo and Otter achieve near-chance accuracy on Bongard-OpenWorld, indicating inability to perform multi-image reasoning OpenFlamingo , Otter IC-129 Sycophancy in VLMs increases with model size: InternVL-1.5-26B (95.8/89.6/86.5%) is more sycophantic than InternVL-1.5-2B (75.6/66.8/98.1%), and InternLM-XComposer2-VL-7B (36.7/28.0/50.7%) more than the 1.8B variant (33.3/20.2/33.0%) InternVL-1.5 , InternLM-XComposer2-VL IC-1290 CLIP, DINO, and DINOv2 as zero-shot natural baselines score below the 50% chance level on Bongard-OpenWorld due to adversarial query selection CLIP / CLIP-ViT (LC) , DINO , DINOv2 IC-1293 The l1 path-norm of PyTorch's pretrained ResNets is approximately 30 orders of magnitude too large for the path-norm generalization bound to be informative on ImageNet-1k ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-130 Amplifying visual-token attention in high layers (16-32) of released VLMs reduces sycophancy while preserving VQA accuracy, indicating that insufficient high-layer visual attention is a key cause of sycophancy LLaVA-1.5 / LLaVA-v1.5 , BLIP-2 , InstructBLIP IC-1307 Released LLMs (CodeLlama 7B/13B/34B, GPT-3.5, GPT-4) achieve limited code-optimization speedups with standard prompting, with the best baseline (GPT-3.5 CoT) reaching only 1.60x versus the 3.66x human reference Code Llama , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1308 Dynamic retrieval-based few-shot prompting substantially improves released LLMs' code optimization, with GPT-4-0613 reaching 76.07% optimization rate and 3.93x speedup (best@8), exceeding the 3.66x human reference Code Llama , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1309 GPT-4-0613 exhibits reduced output diversity relative to GPT-3.5: it outperforms on best@1 but underperforms on best@8 under CoT prompting GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-131 GPT-4 Turbo, GPT-3.5 Turbo, Llama3-8B, Qwen-7B, and iFlytekSpark-13B over-rely on the strong reminder 'the answer is' in prompts as a shortcut, with accuracy dropping sharply when the cue is a random answer rather than the ground truth GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 3 , Qwen , iFlytekSpark-13B IC-1310 Chain-of-thought prompting provides notable code-optimization gains only for larger models (CodeLlama 34B, GPT-3.5, GPT-4) but not for CodeLlama 7B or 13B, consistent with an emergent capability Code Llama , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1317 Llama-2 and Pythia models contain linear representations of space and time that improve with depth and model scale Llama 2 / Llama 2 base , Pythia IC-1318 Individual space and time neurons in Llama-2-7B causally contribute to spatial and temporal predictions Llama 2 / Llama 2 base IC-1319 Larger LLaMA and LLaMA2 models show better calibration on phrase-level tasks but not consistently on sentence- and paragraph-level tasks LLaMA , Llama 2 / Llama 2 base IC-132 GPT-4 Turbo, GPT-3.5 Turbo, Llama3-8B, Qwen-7B, and iFlytekSpark-13B trust authority roles (teacher/judge) more than peer roles (classmate/lawyer) when the cue information is the correct answer GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 3 , Qwen , iFlytekSpark-13B IC-1320 GPT-2 XL (1.5B) exhibits better calibration than larger models from the LLaMA, LLaMA2, and GPT-J families despite having fewer parameters GPT-2 , GPT-J , LLaMA , Llama 2 / Llama 2 base , Vicuna IC-1321 Vicuna-13B, instruction-tuned from LLaMA-13B on user conversations, exhibits worse calibration than its base model LLaMA-13B Vicuna , LLaMA IC-1327 GPT-4, GPT-3.5, Llama2, and Vicuna models underperform human annotators on multistep soft reasoning in natural language narratives, with smaller models scoring near random chance GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base , Vicuna IC-1328 GPT-4's performance on MUSR depends on the type of neurosymbolic scaffolding: program-aided decomposition helps on structured optimization but symbolic belief tracking fails on natural language theory-of-mind GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-133 GPT-4, Claude 3, and Gemini 1.0 Pro do not exhibit detectable watermarks from the red-green, fixed-sampling, or cache-augmented families under black-box statistical tests GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 , Gemini 1.0 Pro IC-1330 All 20 evaluated LLMs improve in multi-turn task-solving with additional tool-use turns and GPT-4-simulated language feedback GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude Instant 1 , Chat-bison-001 , Llama 2 / Llama 2 base , Vicuna , CodeLlama-13B , CodeLlama-34B , Lemur-v1-70B / Lemur-70B-Chat-V1 IC-1331 SIFT and RLHF variants of CodeLlama and Llama-2 perform worse than their base counterparts in multi-turn interaction CodeLlama-13B , CodeLlama-34B , Llama 2 / Llama 2 base , Vicuna , Lemur-v1-70B / Lemur-70B-Chat-V1 IC-1332 Vicuna-v1.5 and CodeLlama-34b-instruct produce format-breaking artifacts (escaped underscores, [python] tags) in 30-100% of code instances due to training data contamination Vicuna , CodeLlama-34B IC-1336 MobileNetV2 (PyTorch pre-trained on ImageNet) exhibits a failure mode under global unstructured L1 pruning at the Pareto-optimal point, with its high kurtosis of kurtoses (64.40) causing very-low-magnitude layers to be entirely pruned and disconnect the network MobileNetV2 , VGG / VGG13 , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-1337 InstructPix2Pix is more effective at editing color than at preserving category in visual concept editing InstructPix2Pix IC-1338 Chinchilla 70B and Llama 2 7B, trained primarily on text, compress ImageNet patches and Librispeech audio better than domain-specific compressors PNG and FLAC Chinchilla , Llama 2 / Llama 2 base IC-1339 Chinchilla 1B's compression rate improves with increasing sequence length across text, image, and audio, demonstrating in-context learning without gradient updates Chinchilla IC-134 Mistral Large V2 as a judge on SummEval coherence systematically avoids extreme ratings (1 and 5), while human annotators assign over 24% of items a median rating of 5 Mistral Large V2 IC-1340 Chinchilla 70B produces coherent autoregressive continuations of text, image, and audio data when used as a compressor, outperforming gzip in sample quality Chinchilla IC-1341 DRUM's standard datalog rule extraction is fundamentally incomplete because its predictions depend on counting distinct rule-body matches DRUM IC-1342 Assigning socio-demographic personas to LLMs causes significant reasoning performance degradation across all four models studied, manifesting as both explicit abstentions and implicit reasoning errors GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 2 / Llama 2 base IC-1343 Task-agnostic de-biasing prompts are ineffective at reducing persona-induced reasoning bias in ChatGPT-3.5, while task-dependent expertise prompts are effective but lack generalizability GPT-3.5 / ChatGPT-3.5 IC-135 Linear relational embeddings for factual relations form in OLMo-7B, OLMo-1B, and GPT-J when subject-object co-occurrence frequency exceeds model-specific thresholds, with r=0.82 correlation between log co-occurrence and causality across all pretraining stages OLMo / OLMo base , GPT-J IC-1355 Instruction-tuned VLMs fail to follow multiple-choice format in reasoning questions, with InstructBLIP frequently returning blank responses InstructBLIP , LLaVA , LLaMA-Adapter v2 , mPLUG-Owl , Otter IC-1356 GPT-3.5 turbo achieves over 90% agreement with human judgments when evaluating VLM responses on open-set questions GPT-3.5 / ChatGPT-3.5 IC-136 LRE quality metrics from OLMo-7B predict pretraining term frequencies in GPT-J (trained on different data) with approximately 70% within-magnitude accuracy for object frequencies, outperforming log-probability-only features by about 30% OLMo / OLMo base , GPT-J IC-1361 GPT-4 and other state-of-the-art LLMs achieve near-human accuracy in inferring personal attributes from unstructured text GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base , Claude Instant 1 , PaLM 2 IC-1362 State-of-the-art text anonymization is insufficient to prevent GPT-4 from inferring personal attributes GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1363 Current model alignment does not filter privacy-invasive prompts across major LLM providers Llama 2 / Llama 2 base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , PaLM 2 IC-1364 GPT-4 can extract personal information from users through adversarial chatbot conversations GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1369 Successor heads that increment ordinal-sequence tokens exist in Pythia, GPT-2, and Llama-2 models from 31M to 12B parameters Pythia , GPT-2 , Llama 2 / Llama 2 base IC-137 Pre-trained ResNet34 and ViT-B features on CIFAR-100 exhibit a block-diagonal class-correlation structure, with ViT-B showing higher intra-class correlation (0.35) than ResNet34 (0.25) ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , ViT IC-1370 MLP0 representations of ordinal-sequence tokens in Pythia-1.4b contain linearly decodable mod-10 features that are causally important for incrementation Pythia , GPT-2 IC-1371 The Pythia-1.4b successor head l12h0 exhibits interpretable polysemanticity, performing successorship, acronym prediction, copying, and greater-than behaviors on natural language data Pythia IC-1372 Successor heads in Pythia-1.4b exhibit a greater-than bias: the OV circuit assigns systematically higher logits to tokens with greater ordinal values than the input, impairing decrementation Pythia IC-1373 All five evaluated LLMs show a strong positional bias in constrained text generation, with first-position constraints nearly always satisfied but last- and arbitrary-position constraints causing major performance drops GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , PaLM 2 , Vicuna IC-1374 Counting difficulty in constrained generation increases with text level and constraint strictness, with exact sentence-level character counts being the hardest condition for all five models GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , PaLM 2 , Vicuna IC-1375 GPT-4's constraint satisfaction improves by approximately 20% after one round of automated feedback but plateaus at 66% even after three additional rounds GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-138 Trojan backdoored Llama-2-7B models and Vicuna-7B-v1.5 exhibit the probe concatenate effect, where concatenating a triggered or jailbroken sample with a harmful probe significantly shifts the model's output distribution away from refusal Vicuna IC-1380 Pythia and OPT small models exhibit non-trivial performance on BigBench tasks that is invisible under beam search but revealed by extensive random sampling Pythia , OPT IC-1386 Fact recall in OPT and LLaMA models degrades by more than 5% relative accuracy when more than 30% of weights are pruned, and similarly when moving from the 30B to the 13B dense model OPT , LLaMA , Pythia IC-1387 In-context learning capabilities in OPT and LLaMA models remain within 5% of dense-model accuracy even at 60-70% sparsity, and show less than 2% difference between the 30B and 1.3B dense OPT models OPT , LLaMA , Pythia IC-1388 In LLaMA-13B, feed-forward layers are more critical than attention layers for fact recall, while both are equally important for in-context learning LLaMA IC-1389 Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMA CLIPBERT , UniVL , VideoCLIP , CLIP4Clip , VioLET , X-CLIP , UniPerceiver , Merlot Reserve , VindLU , InternVideo , mPLUG-2 , Otter , Video-LLaMA , CLIP / CLIP-ViT (LC) , BLIP-2 , GPT-2 , OPT IC-139 3D Gaussian Splatting and its variants (Scaffold-GS, Mip-Splatting) are vulnerable to computation cost attacks via data poisoning, with peak GPU memory increasing up to 21.93x and training time up to 4.97x under unconstrained perturbation 3D Gaussian Splatting , Scaffold-GS , Mip-Splatting IC-1390 Proficiency tests reveal that a substantial portion of correct main-test predictions by VidLMs and ILMs are spurious rather than reflecting robust understanding CLIPBERT , UniVL , VideoCLIP , CLIP4Clip , VioLET , X-CLIP , UniPerceiver , Merlot Reserve , VindLU , InternVideo , mPLUG-2 , Otter , Video-LLaMA , CLIP / CLIP-ViT (LC) , BLIP-2 , GPT-2 , OPT IC-1391 SLD concept removal variants and SD with negative prompts are bypassable by Ring-a-Bell adversarial prompts, increasing attack success rate from single digits to 90-100% for nudity SLD-max , SLD-strong , SLD-medium , Stable Diffusion IC-1392 Title reproduction shows no contamination signal while tag reproduction shows a negative association with GitHub presence and a moderating difficulty effect for GPT-4 and GPT-3.5-turbo GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-1393 Most mainstream LLMs generate value-violating content at high rates (APV 65-80%) across 2,397 morally ambiguous prompts, indicating substantial ethical misalignment Llama 2 / Llama 2 base , LLaMA , ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Falcon , Vicuna , Guanaco , Baichuan , Baichuan 2 , Baichuan2-13B , ChatGLM-6B / ChatGLM-6b-2 , GPT-3 / GPT base IC-1394 ChatGPT demonstrates better ethical value conformity than GPT-4 across multiple prompt generation sources ChatGPT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1395 ChatGPT's ethical violation rate decreases from 70.07 to 57.58 APV when given targeted in-context value instructions generated by VILMO, outperforming baseline alignment methods ChatGPT , Llama 2 / Llama 2 base , GPT-3 / GPT base IC-1399 Invariant GNNs (l=0) consistently fail to distinguish k-hop identical but globally distinct geometric graphs on the k-chain task, regardless of model depth SchNet , DimeNet++ , SphereNet , CoMEt , MACE , GVP , EGNN , CloFNet , ESCN , EquiformerV2 IC-1400 When steerable feature dimension is held constant, increasing the type-l of steerable features does not improve performance of ESCN or EquiformerV2 on IS2RE and S2EF molecular property prediction ESCN , EquiformerV2 IC-1401 GPT-4 serves as a proxy for human judgment on SOTOPIA-EVAL, with strong correlations on goal, financial, and relationship dimensions for model outputs GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1402 Llama-2-70B-Chat underperforms GPT-3.5 across all SOTOPIA dimensions in interactive social scenarios, diverging from static benchmark rankings Llama 2 / Llama 2 base , GPT-3.5 / ChatGPT-3.5 IC-1403 All four evaluated LLMs produce negative scores on social rules and secret-keeping dimensions in SOTOPIA interactions GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base , MPT IC-1404 On SOTOPIA-HARD, GPT-4 achieves significantly lower goal completion than humans and exhibits non-strategic negotiation and excessive compromise behaviors GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1405 OpenFlamingo and Idefics models hallucinate objects not present in images, and increasing ICL shots beyond 4 amplifies hallucinations OpenFlamingo , Idefics IC-1406 OpenFlamingo and Idefics models rarely abstain from answering unanswerable questions, but ICL significantly improves abstention F1 OpenFlamingo , Idefics IC-1407 OpenFlamingo and Idefics models perform near random chance on compositional image-text matching, and ICL has almost no effect on atomic foils OpenFlamingo , Idefics IC-1408 OpenFlamingo and Idefics models generate low-quality explanations in zero-shot, but ICL and model scale significantly improve explanation CIDEr OpenFlamingo , Idefics IC-1409 FF blocks in BERT and GPT-2 modify token-to-token contextualization, with the effect concentrated in specific layers and targeting specific linguistic compositions rather than simple word co-occurrence BERT , MultiBERTs , RoBERTa / RoBERTa-L , GPT-2 , OPT IC-1410 FF's contextualization effects in BERT and GPT-2 are largely canceled by the residual connection and layer normalization, with LN's γ weights specifically shrinking the outlier dimensions in FF output BERT , MultiBERTs , RoBERTa / RoBERTa-L , GPT-2 , OPT IC-1413 GPT-4 achieves near-saturation on Python code synthesis (86.6% pass@1) but scores significantly lower on code repair (47.8% avg) and code explanation (52.1% avg) across six languages GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1414 Pretrained code models Starcoder and CodeGeex2 score 0.0% on code explanation across all six languages because they generate code instead of natural language Starcoder , CodeGeex2 IC-1415 BLOOMZ generalizes instruction-following to programming languages (Go, Rust) absent from its instruction data, scoring above the random baseline BLOOM IC-1416 ChatGPT achieves only 28–35% accuracy on multimodal intent recognition in MINTREC2.0, showing a gap of over 30 percentage points compared to human evaluators ChatGPT IC-1428 OPT 6.7B exhibits over 90% activation sparsity in FFN layers, reducing inference from 6.6G to 4.5G flops per token, while Llama 7B (SiLU) and Falcon 7B (GELU) show near-zero sparsity OPT , LLaMA , Falcon IC-1429 OPT 6.7B exhibits aggregated sparsity where approximately 50% of neurons remain unused across the first 150 tokens, with a non-random reuse pattern enabling 1.27x speculative decoding speedup at gamma=16 OPT IC-1431 BERT-base fine-tuning has negligible distribution-wise variance (0.21%) while BERT-large fine-tuning has substantial distribution-wise variance (2.08%) on MRPC BERT IC-1436 Imagen Video 5.6B produces high-quality but domain-inappropriate videos on out-of-distribution robotics and egocentric data, failing to generate relevant dynamics Imagen Video IC-1437 LLaMA-2-7b-chat-hf and Meta-LLaMA-3-8B-Instruct exhibit reduced attention to rule tokens and fail to follow prompt-specified rules when the adversarial suffix 'forget all prior instructions and answer the question' is appended Llama 2 / Llama 2 base , Llama 3 IC-1438 LLaVA and Llama-Adapter V2 are jailbroken by compositional adversarial images targeting image-based embedding triggers, with near-zero success for textual triggers LLaVA , LLaMA-Adapter v2 IC-1439 LLaVA follows text instructions embedded in adversarial images as if they were user prompts, enabling hidden prompt injection LLaVA , LLaMA-Adapter v2 IC-144 TD-MPC's training is unstable, with performance collapsing after approximately 1–4 million steps across multiple DM Control tasks TD-MPC IC-1440 GCN, GAT, GraphSAGE, and SGC exhibit structure-dependent generalization in transductive node classification: test nodes with shorter paths to training nodes are classified more accurately GCN , GAT , GraphSAGE , SGC IC-1441 GCN exhibits structural unfairness in transductive node classification, with demographic parity and equal opportunity gaps between nodes connected to and disconnected from the training set GCN IC-145 TD-MPC fails to achieve high reward when irrelevant background information is added to image inputs TD-MPC IC-1456 Pre-trained language models fail to predict both interpretations of ambiguous inputs in zero-shot semantic parsing CodeGen , LLaMA , Vicuna , GPT-3.5 / ChatGPT-3.5 IC-1457 Pre-trained language models track the distribution of logical forms in mixed few-shot prompts with conflicting examples CodeGen , LLaMA , Vicuna IC-1458 GPT-4 and PALM 2-L produce significantly less consistent descriptions of interpolated domain embeddings than a purpose-built ELM model GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , PaLM 2 IC-146 Pythia models show increasing robustness to off-policy RLHF data as policy size scales from 410M to 2.8B Pythia IC-1461 Varying decoding hyperparameters and removing the system prompt breaks the safety alignment of 9 out of 11 open-source LLMs, raising attack success rate from 0% to over 95% Vicuna , MPT , Falcon , Llama 2 / Llama 2 base IC-1462 GPT-3.5-turbo is substantially more robust to the generation exploitation attack, with attack success rate of only 7% compared to over 95% for open-source models GPT-3.5 / ChatGPT-3.5 IC-1463 ChatGPT achieves 34.9% average F1 on zero-shot NER across 43 datasets spanning 9 domains ChatGPT IC-1464 Vicuna-7B and Vicuna-13B achieve only 14.2% and 18.0% average F1 on zero-shot NER, trailing ChatGPT by over 20 points Vicuna IC-1465 InstructUIE-11B achieves 81.16% average F1 on 20 in-domain NER datasets and 49.4% on out-of-domain evaluation InstructUIE IC-1467 DECAF achieves PVE 9.65 and f-score 89.6 on the DECAF validation set with a runtime of 19.59 seconds per image DECAF IC-1468 DECAF's optimization-based fitting degrades under significant self-occlusion where the hand covers more than half the face DECAF IC-1469 Progress on standard ImageNet generalization benchmarks is 2.5x faster than progress on crowdsourced global data (DollarStreet, GEODE) across 98 vision models CLIP / CLIP-ViT (LC) , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , DINOv2 , ViT , ConvNeXt , RegNet , MobileNetV3 , VGG / VGG13 , HRNet , FLAVA , EVA-CLIP , MLP-Mixer , EdgeNeXt , ReXNet IC-147 CLIP's OOD performance on rendition domains is largely an artifact of domain contamination in its web-scale training data CLIP / CLIP-ViT (LC) IC-1470 Geographic disparities (Europe-Africa accuracy gap) are large across all 98 models and have more than tripled between least and best performing models on DollarStreet CLIP / CLIP-ViT (LC) , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , DINOv2 , ViT , ConvNeXt , RegNet , MobileNetV3 , VGG / VGG13 , HRNet , FLAVA , EVA-CLIP , MLP-Mixer , EdgeNeXt , ReXNet IC-1471 Common robustness interventions (AugMix, CutMix, Deep AugMix, texture debiasing, antialiasing) and scaling of data or model size do not resolve geographic disparities in released vision models ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , CLIP / CLIP-ViT (LC) IC-1472 DINOv2 (86M parameters) achieves the smallest GEODE geographic disparity (2.46% Europe-Africa gap) among all 98 models in the testbed DINOv2 IC-1477 TAPe performs poorly compared to a simple CNN for protein fitness prediction TAPe IC-148 Language models represent semantically equivalent inputs from different data types (languages, code, images, audio) close together in intermediate layers, with the shared space scaffolded by the model's dominant language Llama 2 / Llama 2 base , Llama 3 , Baichuan 2 , BLOOM , LLaVA , Chameleon , SALMONN IC-1484 GP-UNIT's FID degrades under noisy inputs in reference-guided mode but paradoxically improves in latent-guided mode GP-UNIT IC-1485 Sketch Transformer's FID degrades from 31.49 to 404.01 under Gaussian noise at the highest tested intensity Sketch Transformer IC-1486 HiFaceGAN's FID degrades from 34.83 to 320.41 under Gaussian noise at the highest tested intensity for face super-resolution HiFaceGAN IC-1487 CycleGAN's FID degrades from 76.92 to 180.82 under Gaussian noise for horse-to-zebra translation CycleGAN IC-1489 State-of-the-art foundation models (CLIP, GPT-3.5-turbo, and others) score well below elementary students on multimodal K-12 STEM questions CLIP / CLIP-ViT (LC) , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , ViLBERT , 12-in-1 , UNITER , ViRTex , UnifiedQA , GloVe IC-149 Intervening in the shared representation space using the dominant language (English) predictably changes model outputs for other data types, demonstrating the space is causally used rather than a vestigial byproduct Llama 3 , Llama 2 / Llama 2 base , Chameleon , SALMONN IC-1490 Zero-shot CLIP is overconfident on STEM questions, with softmax confidence loosely related to actual accuracy CLIP / CLIP-ViT (LC) IC-1491 CLIP zero-shot performance on STEM saturates across model sizes, with only 3.6 points of variation from smallest to largest variant CLIP / CLIP-ViT (LC) IC-1496 A single frozen transformer block from LLaMA-7B consistently improves performance across diverse visual tasks when appended to existing visual encoders LLaMA IC-1497 LLaMA-7B's frozen transformer block amplifies informative visual tokens, producing feature activations with higher Miou against ground-truth segmentation masks than both the baseline ViT and the model's own attention scores LLaMA IC-1498 The benefit of frozen LLM transformer blocks for visual encoding is scale-dependent: OPT blocks below 1.3B parameters degrade ViT-s performance while blocks at 1.3B and above improve it OPT IC-1499 Self-rationalization quality and task accuracy scale with model size across GPT-3, FLAN-T5, and LLaMA on five QA datasets GPT-3 / GPT base , FLAN-T5 , LLaMA IC-150 Second-order effects of CLIP's MLP neurons are concentrated in late layers (8–10 of 12 in ViT-B/32) CLIP / CLIP-ViT (LC) IC-1508 LLMs with in-context learning translate Kalamang-English at 44.7/45.8 CHRF, falling short of the human baseline of 51.6/57.0 CHRF GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , GPT-3 / GPT base , Llama 2 / Llama 2 base , LLaMA IC-1509 Kalamang-English translation performance on MTOb increases with model size within the Llama and Llama 2 families, and GPT-4 outperforms Text-davinci-003 LLaMA , Llama 2 / Llama 2 base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3 / GPT base IC-151 Each CLIP neuron's second-order effect is approximately a single linear direction in the joint text-image space, significant for fewer than 2% of images CLIP / CLIP-ViT (LC) IC-1510 Without retrieved context, LLMs are unable to translate Kalamang, and among context types, retrieved parallel sentences are most beneficial, followed by word list entries, then grammar book passages GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , GPT-3 / GPT base , Llama 2 / Llama 2 base , LLaMA IC-1518 Domain finetuning of LLaMA 2 7B, LLaMA 2 13B, and GPT-2 XL on PubMed causes topic and style priors to shift dramatically, accounting for the majority of the probability change, while factual knowledge learning contributes only a small fraction GPT-2 , Llama 2 / Llama 2 base IC-1519 Topic and style biases in LLaMA 2 7B are learned like simple features (rapidly, with minimal capacity, concentrated at the first few tokens, magnified by learning rate) while factual knowledge is learned like complex features (slowly, requiring significant capacity, uniformly across positions, unaffected by learning rate) Llama 2 / Llama 2 base IC-152 CLIP's polysemantic neurons encode spurious correlations between unrelated concepts that can be exploited to generate adversarial misclassifications CLIP / CLIP-ViT (LC) IC-1520 OpenCLIP's per-sample zero-shot accuracy on ImageNet-based OOD benchmarks is strongly correlated with the perceptual similarity between that sample and its nearest neighbor in LAION-400M OpenCLIP IC-1523 Five AI assistants (Claude-1.3, Claude-2.0, GPT-3.5-turbo, GPT-4, Llama-2-70B-Chat) consistently exhibit sycophancy across four varied free-form text-generation tasks Claude 1.3 , Claude 2.0 , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 2 / Llama 2 base IC-153 ResNet50 relies on flower petals and green background features as shortcuts when classifying bee images ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-154 CLIP ViT-L/14 text embeddings fail to capture fine-grained visual class similarities, ranking rottweiler and doberman at position 828 behind unrelated pairs CLIP / CLIP-ViT (LC) IC-1540 CLS-token attention maps in pretrained ViT-t/16 exhibit high inter-layer correlation (cosine similarity up to 0.97) concentrated in layers 3–10, and MSA block outputs show high CKA in layers 2–8 ViT IC-1543 VGG19, ResNet50, ViT-Base, and DeiT-Base (ImageNet pretrained) achieve near-zero accuracy under query-based black-box attacks with 1000–10000 queries VGG / VGG13 , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , ViT , DeiT IC-1544 The latent spaces of pretrained foundational models across vision and text are not related by a single class of geometric transformations; the optimal alignment depends on the specific model pair, architecture, and dataset. CLIP / CLIP-ViT (LC) , ViT , RexNet 100 , BERT-base-cased , BERT , ELECTRA-base-discriminator , RoBERTa / RoBERTa-L , ALBERT-base-v2 , XLM-RoBERTa-base IC-1549 All 28 evaluated LMs exhibit gender bias on non-stereotypical sentence pairs, with fairness scores between 9% and 41% Pythia , GPT-J , OPT , Llama 2 / Llama 2 base , MPT , OLMo / OLMo base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-155 All 13 evaluated MLLMs perform at or near random guessing on MediConfusion, with confusion scores often exceeding 90%, indicating they cannot distinguish visually dissimilar radiology image pairs LLaVA , BLIP-2 , InstructBLIP , DeepSeek-VL2 , Molmo , LLaVA-Med , RadFM , Med-Flamingo , GPT-4o , O1 / OpenAI-o1-preview , Claude 3 , Gemini 1.5 / Gemini Pro 1.5 , Gemini IC-1550 All evaluated LMs systematically prefer male pronoun completions in the non-stereotypical portions of Winobias and Winogender, with margins exceeding 40% Pythia , GPT-J , OPT , Llama 2 / Llama 2 base , MPT , OLMo / OLMo base , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-1551 No consistent relationship between model size and gender fairness scores is observed across six LM families Pythia , OPT , Llama 2 / Llama 2 base , MPT , OLMo / OLMo base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-1552 Deduplication of pretraining data does not consistently improve gender fairness in Pythia models Pythia IC-1553 GPT-J, GPT-2-XL, and Llama-13B decode approximately 48% of tested relations via a linear transformation on the subject representation, and this structure causally influences predictions GPT-J , GPT-2 , LLaMA IC-1554 LRE faithfulness in GPT-J is concentrated in intermediate layers and drops sharply in later layers, consistent with a mode switch from relational encoding to next-token prediction GPT-J IC-1555 GPT-J's internal representations contain correct factual knowledge even when the model outputs falsehoods under repetition or instruction distraction prompts GPT-J IC-1558 Code LLaMA 13B maintains 99.4% passkey retrieval at 128k context despite perplexity rising from 2.37 to 2.54 between 98304 and 131072 tokens CodeLlama-13B IC-156 Gemini models show substantially lower confusion scores than other MLLMs yet still perform at or near random guessing, suggesting their bottleneck is medical knowledge or reasoning rather than visual encoding Gemini 1.5 / Gemini Pro 1.5 , Gemini IC-157 GPT-4o's MediConfusion performance is robust to prompt format while InstructBLIP is highly sensitive and LLaVA-Med fails completely on multiple-choice evaluation GPT-4o , InstructBLIP , LLaVA-Med IC-1576 Base and aligned LLMs share 77.7% of top-1 token predictions, with distribution shifts concentrated in stylistic tokens rather than knowledge content Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Vicuna IC-1577 Base LLMs prompted with URiAL (3 restyled in-context examples + system prompt) match or surpass their SFT/RLHF-aligned counterparts on multi-aspect evaluation Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 2 / Llama 2 base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1578 XLM-R-XL without instruction tuning produces [pad] tokens and fails to complete instruction-following tasks Xlm-R IC-1579 ADM's noise prediction network exhibits exposure bias: during iterative sampling the l2-norm of its ε prediction is systematically larger than during training, and the sampling distribution variance exceeds the training variance with error accumulating toward the end of the chain ADM IC-158 Fine-tuning LLaVA-Med on MediConfusion training pairs cannot achieve 100% training accuracy, indicating the vision encoder's embeddings are fundamentally ambiguous for the confusing pairs LLaVA-Med IC-1580 The official Llama 2-7B checkpoint fails to generate valid numerical responses for 3D-dependent molecular properties, with a valid answer rate of only 23% for SCF energy Llama 2 / Llama 2 base IC-1584 LLaMA-7B and GPT-J-6B fail to interpret textual emphasis markers, with marked prompting degrading performance substantially LLaMA , GPT-J IC-1585 LLaMA-7B and GPT-J-6B exhibit positional bias in instruction following: zero-shot performance varies significantly when the instruction is moved from after to before the context LLaMA , GPT-J IC-1586 In LLaMA-7B, steering all attention heads degrades JSON format accuracy below zero-shot, while steering a subset of 50-100 heads selected via multi-task profiling raises it to 96.64; performance varies dramatically across the 32 layers and individual heads LLaMA IC-159 GPT-4o achieves 55.6% accuracy on creation, 74.8% on math, and 68.1% on code as a preference judge, and is outperformed by domain-specific 7B models on those tasks GPT-4o IC-1598 Retrieval augmentation improves GPT-3.5-turbo-4k on long-context tasks but not GPT-3.5-turbo-16k GPT-3.5 / ChatGPT-3.5 IC-1599 Self-repair at equivalent compute budget provides only modest and inconsistent gains over i.i.d. sampling for CodeLlama-13B-Instruct, GPT-3.5, and GPT-4 on HumanEval and APPS CodeLlama-13B , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-160 Pythia-70m and Gemma-2-2b implement subject-verb agreement across a relative clause via a circuit of number detectors, PP/RC boundary detectors, and verb form promoters, with Gemma-2-2b additionally using NP number trackers Pythia , Gemma 2 IC-1600 Replacing a model's self-generated feedback with a stronger model's feedback consistently improves self-repair beyond both the i.i.d. baseline and the self-repair baseline CodeLlama-13B , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1601 GPT-4's self-generated feedback is significantly less effective than human programmer feedback for code repair, with the gap widening on harder problems GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1602 ResNet18, ResNet34, and MobileNetV2 pre-trained on CIFAR10 have decision functions well-approximated by a kernel machine using the trace NTK, with Kendall-τ correlations of 0.776, 0.786, and 0.700 ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , MobileNetV2 IC-1603 ResNet18's CIFAR10 classification decisions are driven by the bulk of training data rather than a sparse set of exemplars, as revealed by trntk data attribution ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-161 Linear probes on Pythia-70m and Gemma-2-2b trained on the ambiguous Bias in Bios set rely on gender as a spurious feature, with gender accuracy far exceeding profession accuracy Pythia , Gemma 2 IC-1610 Llama-2-7b-chat underperforms on small molecule editing tasks due to limited domain-specific pretraining Llama 2 / Llama 2 base IC-1611 Galactica-6.7b fails on protein secondary structure editing tasks, producing hit ratios below random mutation Galactica-6.7B IC-1612 GPT-3.5-turbo achieves the best and most stable performance across all three drug types in conversational drug editing GPT-3.5 / ChatGPT-3.5 IC-1618 SAM alone has limited generalization for semantic segmentation, producing ambiguous multi-mask outputs without semantic categories SAM , PerSAM IC-1619 DINOv2's patch-level features outperform CLIP and MAE for cross-image semantic feature matching DINOv2 , CLIP / CLIP-ViT (LC) , MAE IC-162 The majority of subject-verb agreement performance in Pythia-70m is explained by approximately 100 SAE feature nodes and in Gemma-2-2b by approximately 500 nodes, compared to approximately 1500 and 50000 neurons respectively Pythia , Gemma 2 IC-163 LLaVA-1.5, LLaVA-Next, and GPT-4V show near-zero accuracy on GUI grounding benchmarks while achieving 50-85 on general image grounding (RefCOCO+), indicating a failure mode specific to GUI grounding scenarios LLaVA-1.5 / LLaVA-v1.5 , LLaVA-NeXT / LLaVA 1.6 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-1630 Llama and Pythia models represent entity-attribute bindings via additive binding id vectors that form a continuous subspace with metric structure LLaMA , Pythia IC-1631 Binding id mechanism fidelity increases with model size in both Llama and Pythia families LLaMA , Pythia IC-1632 Tulu-13B uses a direct binding mechanism rather than binding ids for multiple-choice question tasks Tulu IC-164 Llama-3-8B and Llama-2-7B fail to learn out-of-distribution functions through in-context learning, defaulting to in-distribution predictions Llama 3 , Llama 2 / Llama 2 base IC-165 Llama-3-8B performs algorithm selection during in-context learning, selecting the classification criterion with the lowest test error on ambiguous natural language tasks Llama 3 IC-166 A 1-dimensional subspace in a single layer encodes the context-versus-prior decision in Llama-3.1-8B, Gemma-2 9B, and Mistral-v0.3 7B, and setting this subspace steers the released (non-fine-tuned) models' behavior Llama 3.1 , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-167 Adding a PCA-derived control vector to the middle-layer residual stream improves logit-based reasoning accuracy on Pythia-1.4b, Pythia-2.8b, and Mistral-7B-Instruct Pythia , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-168 Control vectors derived from BABI improve GSM8K accuracy and vice versa on Mistral-7B-Instruct, indicating a task-general reasoning direction in the residual stream Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-169 BT-based, DPO-based reward models, and GPT-4 as judge all exhibit significant length bias, with their scores correlating with output length rather than quality GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Internlm2-Reward , Qwen 2 , Eurus-RM-7B , Llama 3 IC-170 GPT-4, GPT-3.5, and Claude-3.5-Sonnet rely heavily on parametric knowledge in RAG settings, producing ungrounded responses with high answered ratios and low trust-scores GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Claude 3.5 IC-171 ICL prompting produces binary response patterns in released LLMs, with answered ratios collapsing to near 0% or 100% rather than calibrated refusal, making prompting ineffective for RAG groundedness Llama 2 / Llama 2 base , Llama 3 , Llama-3.2-3B , Qwen2.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3.5 IC-174 RAG reduces model abstention and LLMs hallucinate rather than abstain when the retrieved context is insufficient to answer the query Gemini 1.5 / Gemini Pro 1.5 , GPT-4o , Claude 3.5 , Gemma 2 IC-175 Context-sufficiency performance is scale-dependent: larger LLMs achieve high accuracy with sufficient context but still answer correctly 35-62% of the time without it, while smaller models hallucinate or abstain even with sufficient context Gemini 1.5 / Gemini Pro 1.5 , GPT-4o , Claude 3.5 , Gemma 2 , Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-176 LLaMA 3.1 8B Instruct's KGQA accuracy degrades with increasing numbers of retrieved triples, while GPT-4o-mini's accuracy improves, revealing different context-handling capacities Llama 3.1 , GPT-4o IC-177 GPT-4o mini, GPT-4o, and Llama-3-8B all over-rely on incorrect external context, producing wrong answers at high rates when the context conflicts with their internal knowledge GPT-4o , Llama 3 IC-178 Self-guided confidence reasoning (SCR) outperforms rule-based confidence reasoning (RCR) for GPT-4o and GPT-4o mini, but RCR outperforms SCR for Llama-3-8B GPT-4o , Llama 3 IC-179 GPT-4o mini, GPT-4o, and Llama-3-8B all calibrate confidence in their internal answers significantly better than confidence in external contexts GPT-4o , Llama 3 IC-180 GPT-4o-mini's resistance to incorrect context depends on the position of the context relative to the question in the prompt GPT-4o IC-181 Truthfulness in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct is linearly decodable from internal representations at exact answer tokens, with middle-to-late layers being most informative Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-182 Truthfulness encoding in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct is skill-specific rather than universal; probing classifiers do not meaningfully generalize across different task types beyond logit-based baselines Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-183 Error types in Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct are linearly predictable from internal representations, encoding fine-grained information beyond binary correctness Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-184 Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct can internally encode the correct answer while externally generating an incorrect one, with the discrepancy most pronounced for error types where the model shows no external preference for the correct answer Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-185 In Pythia-1B and Amber-7B, the probability of memorizing a training sequence scales log-linearly with both the number of repetitions in the corpus and the z-complexity of the sequence Pythia , Amber-7B IC-186 The memorization status of sequences in Pythia-1B and Amber-7B is stationary throughout training: KL-LD fluctuations are mean-reverting with fixed variance, rejecting a random-walk model with p < 10⁻⁸ Pythia , Amber-7B IC-187 Latent memorized sequences in Pythia-1B and Amber-7B can be recovered by adding random Gaussian noise of magnitude 2×10⁻³ to model weights, while un-memorized and unseen sequences cannot Pythia , Amber-7B IC-188 LLMs show constraint-type-specific performance on system message following, with weaker models exhibiting large variance across constraint categories GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 , Llama 3.1 , Mixtral , GPT-3.5 / ChatGPT-3.5 , Qwen 2.5 72B Instruct , Qwen 2 , GLM-4 , DeepSeek-V2-0628 , Moonshot-v1-8k IC-189 Most LLMs show degraded instruction satisfaction when user instructions conflict with system messages, indicating difficulty in prioritizing system message constraints GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 , Llama 3.1 , Mixtral , GPT-3.5 / ChatGPT-3.5 , Qwen 2.5 72B Instruct , Qwen 2 , GLM-4 , DeepSeek-V2-0628 , Moonshot-v1-8k IC-190 LLMs show progressive degradation in system message constraint following across multi-turn conversations, with dependent conversations degrading faster than parallel ones GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 , Llama 3.1 , Mixtral , GPT-3.5 / ChatGPT-3.5 , Qwen 2.5 72B Instruct , Qwen 2 , GLM-4 , DeepSeek-V2-0628 , Moonshot-v1-8k IC-191 Attention allocated to system messages correlates with following ability, and models do not strictly distinguish system from user messages based on marker tokens GLM-4 , Llama 3.1 , Qwen 2 IC-192 CLIP's contrastive image-text training objective hinders its ability to rank or order images, yielding near-chance performance on ranking tasks in both zero-shot and fine-tuned settings CLIP / CLIP-ViT (LC) IC-193 The OpenCLIP ResNet-50 model trained on CC12M contains an unintentional backdoor from birthday cake images in CC3M, achieving 98.92% attack success rate OpenCLIP IC-194 Temporal modeling in video models drives representational alignment to early visual cortex, while action classification task drives alignment to late brain areas TSM , I3D , SlowFast , MViT V2 , VideoMAE , Uniformer , TimesFormer , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , VGG / VGG13 , ViT , DeiT , X3D , AlexNet , DenseNet / DenseNet-101 , EfficientNet , RegNet , ResNeXt , WideResNet , Inception , RepVGG , Xception , ConvIT IC-195 Transformers achieve high brain alignment in early visual cortex at much shallower network depth than CNNs TSM , I3D , SlowFast , X3D , MViT V2 , VideoMAE , Uniformer , TimesFormer IC-196 Models trained on Something-Something-V2 (which contains no faces) show reduced alignment with face-selective FFA compared to the same models trained on Kinetics-400 TSM IC-197 Model computational complexity (FLOPs) shows a significant negative correlation with brain alignment in high-level brain areas TSM , I3D , SlowFast , MViT V2 , VideoMAE , Uniformer , TimesFormer , X3D IC-198 Safety-aligned LLMs (GPT-4, GPT-3.5, Gemma2-27b, GPT-4o, Gemma2-9b, Qwen2.5-72b, Mistral-7b, Mixtral-8x22b) are vulnerable to natural prompts semantically related to toxic seed prompts, with attack success rates of 82-99% GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , GPT-4o , Gemma 2 , Qwen2.5 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral IC-199 GPT-4o generates natural jailbreak questions from toxic answers without denial, demonstrating an asymmetry in safety training where forward safety (question-to-answer) does not guarantee reverse safety (answer-to-question) GPT-4o IC-200 Personality-related neurons in Llama-3-8B-Instruct are concentrated in the deeper layers of the network Llama 3 IC-201 Activating neuroticism-positive neurons in Llama-3-8B-Instruct causes the largest decline in general capabilities, while activating conscientiousness-positive neurons improves all benchmarks Llama 3 IC-202 All six evaluated LLMs achieve very low accuracy on OpenRCA, with no model solving any three-element root cause query Claude 3.5 , GPT-4o , Gemini 1.5 / Gemini Pro 1.5 , Mistral Large 2 , Command R+ , Llama 3.1 IC-203 Gemini 1.5 Pro's RCA-Agent accuracy drops 68.4% when code execution fails, far exceeding the drops for Claude 3.5 (17.9%) and GPT-4o (15.6%) Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 , GPT-4o , Llama 3.1 IC-204 GPT-4o performs worse with explicit chain-of-thought prompting than with the original prompt on OpenRCA tasks GPT-4o IC-205 Gemma 2's SAE features exhibit depth-dependent organization, with polysemantic features in early layers and persistent, matchable features in later layers Gemma 2 , Llama 3.1 IC-206 GPT-4o-0513 achieves the highest wb-reward mix score (35.7) on WildBench, with a clear three-tier structure among 40 evaluated LLMs GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.5 / Gemini Pro 1.5 , Llama 3 , Claude 3 , Llama 2 / Llama 2 base IC-207 Open LLMs (Llama-3-8B-Inst, Yi-1.5-34B-Chat) show weaker performance on coding and math tasks compared to proprietary models (GPT-4-turbo-0409, Claude 3 Opus) which perform well across all task categories Llama 3 , Yi , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 IC-208 Llama-3-8B-Inst-SimPO does not outperform Llama-3-70B-Inst on WildBench, contrary to its advantage on AlpacaEval-2.0, but performs comparably on information-seeking and creative tasks Llama 3 IC-209 LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferences GPT-3.5 / ChatGPT-3.5 , GPT-4o IC-210 GPT-4o-mini-2024-07-18 does not exhibit authority bias when used as a judge: appending fabricated references to model responses decreases rather than increases its score GPT-4o IC-211 BLOOM-560M employs the same attention head circuit for indirect object identification in both English and Chinese BLOOM IC-212 GPT-2-small and CPM-distilled converge on nearly identical IOI circuits despite being trained independently on English and Chinese GPT-2 , CPM-distilled IC-213 Qwen2-0.5B-Instruct uses English-specific past tense heads and late FFN layers for morphological marking that is absent in Chinese Qwen 2 IC-214 In-context learning of an outlandish sample produces a much diminished keyword-probability-to-priming relationship compared to in-weight gradient learning in Palm-2 PaLM 2 IC-215 DINOv2's zero-shot attention maps focus on irrelevant foreground objects (vehicles, advertisements) rather than scene structure, degrading its VPR recall on challenging datasets DINOv2 IC-216 DINOv2's value (V) facet from self-attention at layer n-1 encodes the most effective local features for VPR re-ranking, outperforming query and key facets, and layer n-1 outperforms the final layer n DINOv2 IC-217 AnyLoc's VLAD aggregation, learned unsupervised on the gallery, fails to generalise to out-of-distribution queries with large time gaps or seasonal changes AnyLoc IC-218 LLaMA3-8B and other LLMs solve arithmetic via a bag of independent heuristic neurons in middle and late MLP layers rather than a robust algorithm Llama 3 , Pythia , GPT-J IC-219 The bag-of-heuristics mechanism in LLaMA3-8B fails on certain arithmetic prompts due to insufficient total logit contribution from heuristic neurons, not due to a lack of associated heuristics Llama 3 IC-220 In Pythia-6.9B, the bag-of-heuristics mechanism emerges gradually during training and is the primary arithmetic mechanism from the earliest checkpoint showing good performance (23k steps) Pythia IC-221 GPT-4o ReAct success rate drops from 47% on synchronous to 11% on asynchronous planning tasks, and all other tested LLMs show equal or worse performance GPT-4o , Gemini 1.5 / Gemini Pro 1.5 , Claude 3 , Qwen 2 , Llama 3.1 , Gemma 2 IC-222 GPT-4o ReAct failures are dominated by rule violations (transition function) and goal misinterpretation, with the balance shifting from goal-dominant in synchronous to transition-dominant in asynchronous settings GPT-4o IC-223 GPT-4o ReAct shows poor recovery from failures in asynchronous settings, with 58.6% of failed runs making little to no progress toward the goal and significantly higher repeated transitions than in synchronous settings GPT-4o IC-224 GPT-4o ReAct cannot incorporate stochastic state changes, with success rate on cutting tasks dropping from 56% to 1% when a 33% chance of a cut item reverting to uncut is introduced GPT-4o IC-225 ESM3 (1.4B) performs zero-shot protein conformation generation with competitive quality across BPTI dynamics, conformation changing pairs, and intrinsically disordered proteins ESM3 IC-226 ESM3 (pre-trained, without fine-tuning) shows significantly degraded conformation generation validity at low sampling temperatures (t < 0.5) ESM3 IC-227 MSA-based conformation generation methods (AlphaFlow, MSA-subsampling) outperform sequence-based methods (EigenFold, STR2STR, ESMFlow) on conformation changing pair generation AlphaFlow , AlphaFold2 , EigenFold , STR2STR , ESMFlow IC-228 ESM3 (1.4B) fails to capture MD ensemble statistics on the ATLAS benchmark, achieving pairwise RMSD correlation of only 0.08 ESM3 IC-229 GIN, GCN, and GAT achieve only random-chance accuracy on WL-separable ε-tree graph pairs, exposing a gap between theoretical expressivity and practical separation GIN , GCN , GAT IC-230 Six LLMs show distinct value preferences on daily-life moral dilemmas, with significant inter-model differences on core values such as truthfulness and fairness GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base , Llama 3 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Claude 3 IC-231 GPT-4-turbo and Claude-3-Haiku show inconsistent adherence to their providers' stated design principles when facing value conflicts in daily-life dilemmas GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 IC-232 System prompts cannot effectively steer GPT-4-turbo's value preferences in moral dilemmas GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-233 Llama-3-70B instruct model differs from its base model in emotion preferences but not in cultural preferences, indicating post-training (RLHF) shapes emotional values Llama 3 , Llama 2 / Llama 2 base , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-234 SpeechGPT exhibits poor speech-text alignment (ASR-WER 45.00) and degraded response quality in speech-to-speech interaction SpeechGPT IC-235 Salmonn and Qwen2-Audio produce responses containing formatted content and redundant explanations that are unsuitable for speech interaction SALMONN , Qwen2-Audio IC-236 Zero-shot Grounding-DINO and Florence-2 show a significant performance gap on referring expression comprehension compared to their fine-tuned versions Grounding DINO , Florence-2 IC-237 A zero-shot Llama 3 8B, when prompted to select the best bounding box from VLM candidates without fine-tuning, produces results nearly identical to the VLM alone Llama 3 , Grounding DINO , Florence-2 IC-238 LLMs fail to follow user preferences in zero-shot settings, with accuracy below 10% at 10 turns and near zero at 300 turns Claude 3 , Claude 3.5 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Llama 3 , GPT-4o , GPT-4.1 , O1 / OpenAI-o1-preview , O4-mini , O3 , Gemini 1.5 / Gemini Pro 1.5 IC-239 Implicit preference forms (choice-based and persona-driven) are significantly harder for LLMs to follow than explicit preferences at the same context length Claude 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Llama 3 IC-240 Introducing multiple preferences (including conflicting ones) in a conversation improves LLM adherence to the original preference Claude 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-241 Preference following degrades when the preference is placed in the middle of a long conversation, extending the lost-in-the-middle effect to preference tracking Claude 3 IC-242 Most LLMs exhibit higher bias ratios in multi-turn dialogues than in single-turn, with bias accumulating across successive turns Llama 2 / Llama 2 base , Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma IC-243 Bias ratio on certain multi-turn fairness tasks decreases with model size in the Gemma-2 and Qwen2.5 families Gemma 2 , Qwen2.5 IC-244 No LLM demonstrates consistently strong fairness across both comprehension-focused and bias-resistance multi-turn tasks; models show complementary failure patterns Llama 2 / Llama 2 base , Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma IC-245 Pretrained LLMs produce duration-dependent outputs that are incompatible with a discrete token interpretation Llama 3 , Llama 2 / Llama 2 base , Phi-3 , Gemma , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-246 Pretrained LLMs assign coherent semantic meaning to linear interpolations between token embeddings, extending the linear embedding hypothesis to the output space Llama 3 , Llama 2 / Llama 2 base , Phi-3 , Gemma , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-247 Pretrained LLMs are invariant to positional shifts but sensitive to duration scaling of the input Llama 3 , Llama 2 / Llama 2 base , Phi-3 , Gemma , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , GPT-2 IC-248 Instruction fine-tuning causes context reliance under knowledge conflicts to initially increase then decrease (context-parametric inversion) in Llama2-7B, Pythia-6.9B, and Mistral-7B Llama 2 / Llama 2 base , Pythia , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-249 The GitHub data-refined LLC identifies the induction circuit heads in Pythia-70m by distinguishing previous-token and induction heads from other head types across layers 2 and 3 Pythia IC-250 Existing multimodal embedding models show highly uneven performance across MMEB's four meta-task categories, with VQA scores as low as 4.2 and overall scores ranging from 13.3 to 44.7 CLIP / CLIP-ViT (LC) , BLIP-2 , SigLIP , OpenCLIP , UniIR , MagicLens , E5-V IC-251 CLIP's overall MMEB performance drops by 29.4% when task-specific instructions are prepended to queries, with classification degrading by 59.3% CLIP / CLIP-ViT (LC) IC-252 GPT models produce harmful gender stereotypes at higher rates when user names imply a demographic group, with GPT-3.5 Turbo showing the highest rates and open-ended generation tasks most affected GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , O1 / OpenAI-o1-preview IC-253 Post-training reinforcement learning significantly reduces harmful gender stereotypes in GPT models, with the best-fit slope of 0.21 indicating post-RL models have far lower bias than pre-RL versions GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o IC-254 GPT-4o Mini responses to female-sounding names systematically use simpler, more light-hearted, and less technical language compared to male-sounding names across multiple task domains GPT-4o IC-255 GPT-4o, Llama 3.1, and Claude models show varying correlation with human ratings when used as stereotype evaluators, with GPT-4o achieving the strongest gender correlation (ρ=0.86) but weaker racial correlations GPT-4o , Llama 3.1 , Claude 3.5 , Claude 3 IC-256 GPT-2 small's attention product functions p_i^T k^T q p_j are approximately translation-invariant across all 144 heads GPT-2 IC-257 327 DNNs approach or exceed human accuracy on object depth order but are near chance on VPT-basic, while humans show the opposite pattern BEiT , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini , Claude 3 , Stable Diffusion , MAE , DINOv2 , SAM , MiDaS , Depth Anything IC-258 DNN accuracy on 3D perception tasks correlates with ImageNet object classification accuracy, suggesting 3D cues emerge as a byproduct of object recognition training BEiT , Swin Transformer IC-259 Fine-tuned DNNs approach human accuracy on VPT-basic but fail on VPT-strategy, revealing reliance on a brittle feature-based shortcut (object size and location) rather than line-of-sight estimation Swin Transformer IC-260 DeepGate2 achieves an average NDCG@3 of 0.334 and top-10% commonality of 0.226 on QOR prediction across 10 circuit designs DeepGate2 IC-261 DeepGate2 achieves an average F1-score of 0.424 and AUC of 0.804 on logic equivalence identification across 10 circuit designs DeepGate2 IC-262 DeepGate3 achieves an average F1-score of 0.390 and AUC of 0.834 on logic equivalence identification for small circuit designs DeepGate3 IC-263 DeepGate2 and DeepGate3 achieve average SAT solving runtime reductions of 18.24% and 21.31% respectively when used to constrain boolean fence search space DeepGate2 , DeepGate3 IC-264 All 18 evaluated LLMs fail to abstain when the provided context lacks the answer, with performance gaps of 13.6% to 68.4% relative to the original context Phi-3 , Phi-3.5 Mini Instruct , Llama 3 , Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma 2 , GPT-3.5 / ChatGPT-3.5 , GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Command R+ , Claude 3.5 IC-265 Model families show extreme variation in detecting conflicting answers in inconsistent contexts, with phi-3 series at 5.8% average accuracy versus GPT-4 series at 89.35% Phi-3 , Phi-3.5 Mini Instruct , Llama 3 , Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma 2 , GPT-3.5 / ChatGPT-3.5 , GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Command R+ , Claude 3.5 IC-266 GPT-4o drops from 96.3% closed-book accuracy to 47.5% when given counterfactual context that contradicts its parametric knowledge, far below the 95% human accuracy on the same items Phi-3 , Phi-3.5 Mini Instruct , Llama 3 , Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma 2 , GPT-3.5 / ChatGPT-3.5 , GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Command R+ , Claude 3.5 IC-267 Adding a 'conflict' instruction to the prompt degrades GPT-4o and Claude 3.5 Sonnet accuracy on normal (answerable, consistent) contexts by 5% and 2% respectively GPT-4o , Claude 3.5 IC-268 ESM-2 and ProGen-2 zero-shot fitness prediction follows an inverted U-shape as a function of wild type sequence likelihood, with both under- and over-preferred sequences degrading performance ESM-2 , ProGen-2 IC-269 Influence functions on ESM-2 650M reveal a power law tail in training data influence on sequence likelihood, with influence diminishing as Hamming distance from the wild type increases ESM-2 IC-270 Unsupervised finetuning (evo-tuning) on homologous sequences improves ESM-2 650M zero-shot fitness prediction for low-likelihood wild types but harms high-likelihood ones, with optimal threshold at log-likelihood ε = −1.4 ESM-2 , EVE , MSA Transformer , TranceptionEVE , ProGen-2 IC-275 Mistral 7B Instruct exhibits a reasoning-type-dependent failure mode where certain problems are exclusively solvable by one non-deductive reasoning type Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-276 All 14 evaluated VLMs show a large gap between average-case and worst-case accuracy on DynaMath variants, with worst-case at or below 50% of average-case, and the failures are systematic rather than random GPT-4o , Claude 3.5 , Gemini 1.5 / Gemini Pro 1.5 , Qwen2-VL , InternVL2 , LLaVA-NeXT / LLaVA 1.6 , LLaVA-1.5 / LLaVA-v1.5 , DeepSeek-VL , Llama 3.2 IC-277 Open-source VLMs show a clear scaling trend in both average accuracy and reasoning robustness on DynaMath, with larger models performing substantially better Qwen2-VL , InternVL2 , Gemini 1.5 / Gemini Pro 1.5 IC-278 Claude-3.5 Sonnet and GPT-4o exhibit a memorization failure mode, outputting the same answer regardless of visual parameter changes in the problem Claude 3.5 , GPT-4o IC-279 GP-LVM produces less structured latent representations and lower generative-classification accuracy than QEP-LVM on oil flow and MNIST GP-LVM IC-280 CLIP, OpenCLIP, and SigLIP exhibit intra-modal misalignment: intra-modal similarity comparisons are suboptimal for image-to-image and text-to-text retrieval CLIP / CLIP-ViT (LC) , OpenCLIP , SigLIP IC-281 SLIP's intra-modal self-supervised loss reduces intra-modal misalignment, making inter-modal inversion unnecessary for image retrieval SLIP , CLIP / CLIP-ViT (LC) IC-282 GPT-2 XL (1.5B) exhibits lower accuracy but reduced overconfidence (smaller ECE and Brier scores) compared to larger models on the CAT benchmark GPT-2 , Vicuna IC-290 Zeroing out or doubling specific FFN neurons identified by the neuron path method causes significant accuracy changes in ViT and MAE models ViT , MAE-B/16 IC-291 ViT-B/16 and MAE-B/16 exhibit nearly inverted distributions of knowledge neurons across layers despite identical architecture and training data ViT , MAE-B/16 IC-292 Neuron paths in ViT-B/16 show class-specific neuron clustering and semantic similarity between image categories ViT , MAE-B/16 IC-293 ViT-B/16 and ViT-B/32 are largely redundant: retaining only top-5 neurons per layer while zeroing all others preserves most classification accuracy ViT IC-294 Llama-3.1-8B-Instruct with 2-shot prompting achieves limited rationale extraction quality (F1 15.7–48.3) across four text classification datasets Llama 3.1 IC-295 The ViT model's ECE can be reduced to near-zero by trivial mean-replacement recalibration while maintaining test accuracy, but NLL increases from 65.35 to 144.66, demonstrating that ECE and accuracy alone are an insufficient reporting standard for calibration ViT , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , EfficientNet , ConvNeXt IC-296 The degree to which SAE features are active at multiple residual-stream layers increases with model size in Pythia, Gemma 2, Llama 3.2, and GPT-2 Pythia , Gemma 2 , Llama 3.2 , GPT-2 IC-297 Applying tuned-lens transformations to the residual stream decreases the apparent multi-layer SAE feature activity from 54–88% to 37–41% of total variance Pythia IC-298 Weight similarity in open-source LLMs is organized in a depth-dependent structure with adjacent-layer similarity and distinct clusters at specific depths Llama 3.1 , Gemma 2 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-299 Instruction tuning preserves the weight-matrix structure of LLMs, with DOCS scores exceeding 0.7 across all matrices Yi , Llama 3.1 , Gemma 2 IC-300 One MoE expert in Mixtral-8x7B is structurally distinct from the others in many layers Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-301 Yi-1.5-9B-chat exhibits a layer-repetition pattern where a section of layers is duplicated at a later depth Yi IC-302 Llama-3.1 models perform between unigram-inference and bigram-inference on Markov chain ICL, with performance improving monotonically with model scale Llama 3.1 IC-303 Explicitly stating the Markovian structure in the prompt significantly improves Llama-3.1-70B's next-state prediction on the Markov chain task Llama 3.1 IC-304 Instruction-tuned LMs become more vulnerable to prompt-injected data extraction as model size increases from 7B to 70B Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Solar 10.7B , Vicuna , Wizardlm , Qwen1.5 , Platypus2-Instruct-70B IC-305 Mistral-instruct-7b's susceptibility to prompt-injected data extraction follows a U-shaped curve depending on the position of the adversarial prompt within the context window Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-306 Instruction tuning increases the ROUGE score of prompt-injected data extraction by 65.76 on average compared to base models Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-307 56 LLMs on Sorry-Bench show fulfillment rates ranging from below 10% (Claude-2, Gemini-1.5) to above 90% (Mistral-7B-instruct-v0.1, Dolphin-2.6-mixtral-8x7b), with GPT-4o at 30% and Llama-3-70B at 35% GPT-4o , GPT-3.5 / ChatGPT-3.5 , Claude 2.1 , Claude 2.0 , Gemini 1.5 / Gemini Pro 1.5 , Gemini , Llama 3 , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma , Vicuna , OpenChat-3.5-0106 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Zephyr-7B-beta IC-308 Linguistic mutations to unsafe prompts significantly and inconsistently alter safety refusal across models, with persuasion techniques increasing fulfillment by 5-66% and encoding/encryption decreasing it by 15-68% GPT-4o , GPT-3.5 / ChatGPT-3.5 , Llama 3 , Gemma , Vicuna , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , OpenChat-3.5-0106 IC-309 As zero-shot safety judges, GPT-4o achieves 78.9% Cohen's kappa agreement with human annotators while Llama-3-8B-instruct (39.0%) and Mistral-7B-instruct-v0.2 (53.9%) perform substantially worse GPT-4o , GPT-3.5 / ChatGPT-3.5 , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma , Llama Guard 2 , WildGuard , HarmBench Classifier , BERT-base-cased IC-310 Prefilling model responses with 'sure, here is' increases safety fulfillment by 19-58%, and missing prompt template tokens increases fulfillment by 8-30% for Llama-2 and Gemma but not Llama-3 Llama 3 , Llama 2 / Llama 2 base , Gemma IC-311 On siltuximab GRAVY reduction, Lambo-2 achieves the best concept shift while ESM2 produces the most naturalness-disrupted designs Lambo-2 , ESM-2 , WJS IC-312 GPT-4o achieves the highest insight-level Llama-3-eval score (0.60) among all tested LLM backbones on InsightBench multi-step data analytics GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 3 IC-313 All four LLM backbones fail to detect a planted linear trend in incident resolution time when the slope is below 0.1, and detection rates diverge sharply above that threshold GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 3 IC-314 Llama-3-70b as an LLM-based evaluator (Llama-3-eval) produces agent rankings consistent with GPT-4-based G-Eval on InsightBench Llama 3 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-315 CoT prompting (reasoning + instruction) yields larger relative gains for larger LLMs and harder problems in competitive code generation, with the effect reversing for the most capable models Llama 3 , Llama 3.1 , GPT-4o IC-316 Multi-turn code generation without CoT degrades performance for smaller Llama models and GPT-4o compared to single-turn repeated sampling under equal compute budgets Llama 3 , Llama 3.1 , GPT-4o IC-317 More detailed execution feedback (LDB) induces exploitative behavior in Llama 3.1 models, reducing code diversity and hurting performance at large sample budgets Llama 3.1 IC-318 CLIP backbones from different architectures (ViTs and ResNets) trained with the same data and objective exhibit complementary strengths, with an oracle per-image backbone selection improving zero-shot accuracy by up to 43.5% over the best single backbone CLIP / CLIP-ViT (LC) IC-319 Different CLIP backbones exhibit distinct robustness profiles to specific image perturbations, with each architecture being most resilient to a different transformation CLIP / CLIP-ViT (LC) IC-320 Llama-2-13b-chat underperforms Llama-2-7b-chat on fine-grained dimension-level evaluation Llama 2 / Llama 2 base IC-321 GPT-4o selects evaluation dimensions with high precision but low recall, indicating a selective rather than comprehensive strategy GPT-4o IC-322 Past-tense reformulations of harmful requests bypass refusal training in eight released LLMs, while future-tense reformulations are substantially less effective Llama 3 , Claude 3.5 , GPT-3.5 / ChatGPT-3.5 , Gemma 2 , Phi-3 , GPT-4o , R2D2 IC-323 O1-mini and O1-preview reasoning models are vulnerable to past-tense reformulations (84% and 78% ASR) but produce less specific jailbroken outputs than non-reasoning models O1 / OpenAI-o1-preview IC-324 Fine-tuning Gemini Nano 1 on 8 memorization examples causes it to override in-context predictions with in-weight predictions in 2 of 8 cases, while the base model always follows in-context predictions Gemini IC-325 A single FFN-layer weight edit (JailbreakEdit) raises jailbreak success rate to 62–87% on Llama-2-7b-chat, Llama-2-13b-chat, Vicuna-7b, and ChatGLM-6b while preserving safety performance and generation quality on non-triggered queries Llama 2 / Llama 2 base , Vicuna , ChatGLM-6B / ChatGLM-6b-2 IC-326 Jailbreak vulnerability and response style are scale-dependent: Llama-2-13b-chat shows higher post-attack JSR and a shift toward direct compliance (type-5 actions) compared to Llama-2-7b-chat Llama 2 / Llama 2 base IC-327 Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, and several other LLMs produce well-calibrated verbal confidence estimates on classification tasks Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , GPT-4o , Gemma 2 , Mistral-Nemo 12B-Instruct-2407 , Qwen2.5 , Llama-3-2-Vision IC-328 Llama-3-8B-Instruct and Mistral-7B-Instruct-v0.3 are susceptible to confidence-elicitation-guided word substitution attacks, with CEAttack outperforming existing hard-label black-box methods Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-329 GPT-4o is more robust to confidence-elicitation-guided word substitution attacks than open-source LLMs, with lower attack success rates and better confidence calibration GPT-4o IC-330 Low local intrinsic dimension (LIDθ) of the learned manifold predicts memorization in Stable Diffusion v1.5, IDDPM, and StyleGAN2-ADA Stable Diffusion , IDDPM , StyleGAN2-ADA IC-331 In Stable Diffusion v1.5, specific tokens in text prompts drive memorization, and GPT-4-based perturbation of high-attribution tokens reduces SSIM similarity to training images while maintaining CLIP score Stable Diffusion , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , CLIP / CLIP-ViT (LC) IC-332 Llama-3-70B exhibits a friendlier, funnier, and less ethics-focused style than GPT-4 and Claude-3-Opus on Chatbot Arena, and these vibes predict model identity at 80% and user preference at 59% accuracy Llama 3 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3 IC-333 Llama-3-405B overexplains math solutions with structured markdown headings and conversational tone compared to GPT-4o's concise formal notation, achieving 97% model-matching accuracy GPT-4o , Llama 3 IC-334 GPT-4V produces more poetic, emotion-focused image captions compared to Gemini-1.5-Flash's literal descriptions, with 99% model-matching accuracy on COCO GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.5 / Gemini Pro 1.5 IC-335 GPT-4's detection performance as a scoring model is highly sensitive to the prompt, varying from 0.7289 to 0.9682 AUROC, far more than GPT-3.5 or Babbage GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , GPT-3 / GPT base IC-336 Larger proprietary LLMs (GPT-3.5 175B) are more effective universal text detectors than smaller models (Babbage 1.3B, GPT-Neo-2.7B), contradicting prior findings that smaller models are better GPT-3.5 / ChatGPT-3.5 , GPT-3 / GPT base , GPT-Neo IC-337 GPT-3.5-based detection accuracy drops substantially for Russian text (0.8555 AUROC) compared to near-perfect scores for Urdu, Indonesian, and Arabic, suggesting under-training on Russian GPT-3.5 / ChatGPT-3.5 IC-338 Factuality enhancement methods (DoLa, ICD, ITI, TruthX, CD) cause large and consistent declines in context-faithfulness of LLaMA2-7B-Chat and LLaMA2-13B-Chat Llama 2 / Llama 2 base IC-339 Factuality enhancement methods produce inconsistent and modest improvements in factual accuracy on LLaMA2-Chat, with some metrics declining below baseline Llama 2 / Llama 2 base IC-340 GCG jailbreaking attacks exhibit strong model-specific transferability, achieving below 3% ASR on Llama-2-13b-chat and Llama-3.1-8b-instruct but above 90% ASR on Vicuna-13b-v1.5 and Mistral-7b-instruct Llama 2 / Llama 2 base , Llama 3.1 , Vicuna , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , O1 / OpenAI-o1-preview IC-341 The effectiveness of GCG and PAIR attacks on Llama-2-7b-chat is sensitive to the order of adversarial tokens, with swapping the two halves of the GCG suffix reducing the created high-importance region by 23% Llama 2 / Llama 2 base IC-342 Aligned Llama-2-7b-chat allocates 37% perceived-importance to 'bomb' and 21% to 'build' in its intent perception, while unaligned Llama-2-7b shows uniform perceived-importance across all tokens Llama 2 / Llama 2 base IC-343 GPT-4o, Claude-3.5 Sonnet, and GeminiPro-1.5 score below BigDocs-trained open models on BigDocs-Bench tasks requiring long structured code generation GPT-4o , Claude 3.5 , Gemini 1.5 / Gemini Pro 1.5 , Qwen2-VL , Llama 3.2 , Idefics2 IC-344 GPT-4o achieves the highest average score (64.62) on general document benchmarks, outperforming Qwen2-VL-72B (58.40) and GeminiPro-1.5 (57.05) GPT-4o , Qwen2-VL , Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 , Llama 3.2 IC-345 GPT-4o's table2latex outputs lose 63% of the time in human evaluation, with inconsistent formatting (lines, borders, margins) as the primary failure GPT-4o IC-346 Hierarchical and categorical concepts from WordNet are linearly represented in the final-layer space of Gemma-2b and Llama-3-8B, with semantic hierarchy encoded as orthogonality and categorical concepts as polytopes Gemma , Llama 3 IC-347 GPT-3.5-turbo-instruct achieves 53.7% move-matching accuracy on human chess when prompted with PGN notation GPT-3.5 / ChatGPT-3.5 IC-348 Sequential parameter-modifying editing causes progressive degradation of general abilities in GPT-2 XL, Llama-2 7B, and Llama-3 8B, driven by growth in the condition number of the edited matrix GPT-2 , Llama 2 / Llama 2 base , Llama 3 IC-349 Larger LLMs (Llama-2 7B, Llama-3 8B) suffer more severe general ability degradation than smaller models (GPT-2 XL 1.5B) under the same number of sequential edits GPT-2 , Llama 2 / Llama 2 base , Llama 3 IC-350 Editing conceptual knowledge with rome on Llama-2 7B is harder than factual knowledge editing, with the model failing to update concept-instance relationships Llama 2 / Llama 2 base IC-351 GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct show distinct performance profiles across ultra-long, 32k, and 4k context benchmarks GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Qwen 2 , Llama 3.1 , Yi , Llama 3 IC-352 RAG with sufficient retrieved tokens outperforms direct long-context for Qwen2-72b-instruct on >100k tasks, while at 32k the default RAG setting underperforms direct long-context for GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Qwen 2 , Llama 3.1 IC-353 Llama-3.1-instruct 8b and 70b fail the harder NIAH test (sandwich needle) but pass the easier passkey retrieval test Llama 3.1 IC-357 Off-the-shelf foundation models (DINO, CLIP, DINOv2, ViT) exhibit higher variance in their cosine similarity distributions than dataset-specific models, reducing the discriminative power of cosine similarity retrieval DINO , DINOv2 , CLIP / CLIP-ViT (LC) , ViT , CoPlace IC-358 GPT-2-small and Mistral 7B contain circular representations of days of the week and months of the year in their internal activations, discovered via SAE dictionary element clustering GPT-2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-359 Mistral 7B and Llama 3 8B causally use circular subspaces to compute modular arithmetic on days of the week and months of the year Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-360 GPT-2 achieves only trivial accuracy on modular arithmetic tasks for days of the week and months of the year despite containing circular representations GPT-2 IC-361 Mistral 7B's circular representation of days of the week is continuous, mapping intermediate time-of-day values to positions between adjacent weekdays Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-362 LLMs with chain-of-thought prompting predict and simulate human risky choices that are more rational than actual human behavior, correlating more highly with maximum expected value than with human choices Llama 3 , Claude 3 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o IC-363 LLM inferences about others' preferences from observed decisions are highly correlated with human inferences because both assume the decision-maker is rational Llama 3 , Claude 3 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o IC-364 Llama-3-8B hidden states are zero-mean unimodal, with gaussian-like distributions before attention and MLP blocks and laplacian-like distributions in intermediate states Llama 3 IC-365 70B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsity Llama 3 , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-366 In Llama-3-70B, activation sparsifiability varies systematically across depth: Wq/Wk peak in block 0 then decline sharply, Wo peaks at 80-90% mid-model, and Wdown is consistently more sparsifiable than Wgate and Wup Llama 3 IC-367 Sparsifying initial tokens of the prefill phase causes disproportionate degradation in Llama-3-8B due to attention sink behavior Llama 3 IC-368 Larger LMs (GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro) exhibit better calibration than their smaller counterparts (GPT-4o-mini, Claude-3-Haiku, Gemini-1.5-Flash) when verbalizing confidence with certainty phrases GPT-4o , Claude 3.5 , Claude 3 , Gemini 1.5 / Gemini Pro 1.5 IC-369 LMs verbalizing confidence with certainty phrases are better calibrated on SCIQ than on TruthfulQA GPT-4o , Claude 3.5 , Gemini 1.5 / Gemini Pro 1.5 IC-370 Knowledge entropy (sparsity of FFN memory coefficients) decreases consistently during pretraining for OLMo 1B, 7B, and Pythia 1.4B, and this decrease strongly correlates with reduced knowledge acquisition and increased forgetting in continual learning OLMo / OLMo base , Pythia IC-371 Artificially resuscitating inactive memory vectors by scaling the up-projection matrix K improves knowledge acquisition and reduces forgetting, with the effect more pronounced for later-stage OLMo models OLMo / OLMo base IC-372 Language models universally decompose retrieval tasks into request processing in middle layers and entity retrieval in late layers at the last token position GPT-2 , Pythia , Falcon , Llama 2 / Llama 2 base IC-373 In Pythia-2.8B, the specific attention heads and MLPs implementing retrieval depend on superficial input features, and request-patching preserves the natural retrieval mechanism Pythia IC-374 Pythia models are vulnerable to prompt injection via distractor text, and request-patching from a single trusted input restores most of their accuracy Pythia IC-381 Individual knowledge is not parameter-localizable in GPT-J: existing localization methods (KN, ROME, KC) are neither faithful nor reliable GPT-J IC-382 Data commonalities are localizable to a small set of capability neurons in Llama2-7B, Llama2-13B, and GPT-J-6B, and these neurons enhance or degrade performance when manipulated Llama 2 / Llama 2 base , GPT-J IC-383 GPT-4-1106, GPT-3.5-1106, and unfine-tuned CodeLlama-13b achieve 44.3%, 39.5%, and 38.5% API call accuracy respectively on unseen APIs (Level 3) with 3-shot retrieved prompting in the API Pack evaluation GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , CodeLlama-13B IC-384 Llama-3-8B-Instruct shows larger absolute gains from prompt optimization than the stronger Gemma-2-9B-IT, indicating that prompt-optimization benefit is inversely related to base model capability Llama 3 , Gemma 2 IC-385 TAR-bio-v1 retains bio-weaponization knowledge despite appearing to unlearn it; a different prompt template and answer extraction method reveals accuracy above 45% on WMDP-bio Llama 3 IC-386 LLM performance on CS-Bench grows logarithmically with parameter scale within model families Qwen1.5 , Llama 2 / Llama 2 base , Llama 3 , Gemma , InternLM2 , DeepSeek LLM IC-387 OpenAI-o1 models substantially improve CS reasoning over GPT-4o at the cost of 14-30x token consumption GPT-4o IC-388 CS-Bench scores correlate strongly (p > 0.9) with math and code benchmark scores across 12 models Qwen1.5 , Llama 2 / Llama 2 base , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-389 All evaluated LLMs score significantly lower on CS reasoning questions than knowledge questions, with the gap narrowing for stronger models Llama 2 / Llama 2 base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , GPT-3.5 / ChatGPT-3.5 , PaLM 2 , Claude 2.1 IC-390 The parallelotope volume of modality embeddings from LanguageBind, VAST, and Valor on MSR-VTT is strongly correlated with their downstream R@1 retrieval performance VAST , LanguageBind , Valor IC-391 Llama-3 and Qwen-1.5 models exhibit position bias in LM-as-a-judge, retrieval-augmented QA, and math reasoning, with larger models showing less bias Llama 3 , Qwen1.5 , Qwen 2.5 72B Instruct IC-392 Fuyu-8B and GPT-4V exhibit position bias in visual recognition, with model performance depending on where the target object appears in the image Fuyu , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-393 GPT-4-turbo, Llama-3.1-8B-Instruct, and OpenAI Moderation show declining hate speech detection accuracy as sentence implicitness increases, with very low success rates in the highest implicitness ranges GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 3.1 , OpenAI Moderation IC-394 Text generation in SDXL, DeepFloyd IF, and SD3 is controlled by less than 1% of parameters concentrated in specific cross- or joint-attention layers, and these layers are specialised for text content rather than visual template Stable Diffusion , DeepFloyd IF IC-395 Most mainstream LLMs exhibit positive ADCE across five tasks, indicating reliance on deep structure for problem-solving, with ADCE strongly correlated with accuracy (r² > 0.7) Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Mixtral , GPT-3.5 / ChatGPT-3.5 , GPT-4o , Claude 3 , Claude 3.5 IC-396 Closed-source LLMs (GPT, Claude) rely more on deep structure than open-source LLMs (Llama, Mistral), and open-source models' surface sensitivity decreases with model scale Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Mixtral , GPT-3.5 / ChatGPT-3.5 , GPT-4o , Claude 3 , Claude 3.5 IC-397 Mistral 7B Instruct and Llama 3 8B Instruct exhibit systematic misalignment between their operational semantics of subjective phrases and human expectations, producing unexpected side effects when steered with certain phrases Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-398 Ablating a single safety attention head in Llama-2-7b-chat increases attack success rate from 0.04 to 0.64 and in Vicuna-7b-v1.5 from 0.27 to 0.55, by modifying only 0.006% of parameters Llama 2 / Llama 2 base , Vicuna IC-399 Safety attention heads overlap significantly between Llama-2-7b-chat and Vicuna-7b-v1.5, indicating that pre-training shapes safety capability Llama 2 / Llama 2 base , Vicuna IC-400 Safety attention heads function as feature extractors: modifying the attention pattern (Wq/Wk) has far greater safety impact than modifying the value (Wv) in Llama-2-7b-chat Llama 2 / Llama 2 base IC-401 Ablating safety attention heads minimally degrades helpfulness on zero-shot tasks and also impairs course-correction capability in Llama-2-7b-chat Llama 2 / Llama 2 base IC-402 In LLaMA3-8B, LLaMA2-13B, and Mistral-7B, soft-prompt information flow peaks in shallow layers (2–10) and reasoning correctness depends on whether deeper layers redirect attention away from soft prompts to earlier reasoning steps Llama 3 , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-407 Safety-aligned LLMs (Llama-2-chat, Llama-3-instruct, Gemma, GPT-3.5, GPT-4o, R2D2) achieve 100% jailbreak attack success rate under adaptive prompt-and-suffix attacks on 50 harmful requests Llama 2 / Llama 2 base , Llama 3 , Gemma , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , R2D2 IC-408 Claude models (2.0, 2.1, 3 Haiku, 3 Sonnet, 3 Opus, 3.5 Sonnet) achieve 100% jailbreak attack success rate under prefilling attacks via the Anthropic API Claude 2.0 , Claude 2.1 , Claude 3 , Claude 3.5 IC-409 Knowledge editing methods correct verified hallucinations in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B far less effectively than their scores on existing benchmarks suggest Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-410 Knowledge editing can degrade generalization performance below pre-edit levels in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-411 Llama2-7B, Llama3-8B, and Mistral-v0.3-7B do not reason with edited knowledge in multi-hop questions, as editing methods mostly underperform pre-edit portability scores Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-412 Edited knowledge in Llama2-7B is significantly less robust to adversarial prompts than in Llama3-8B and Mistral-v0.3-7B Llama 2 / Llama 2 base , Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-413 Factors of variation in ImageNet-X are linearly decodable from the second-to-last-layer representations of ImageNet-pretrained ResNet50 and ViT-B/16 ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , ViT IC-414 LLaVA-1.5-7B and LLaVA-1.5-13B exhibit severe performance degradation when H2O KV cache compression is applied in multimodal settings LLaVA-1.5 / LLaVA-v1.5 IC-415 VLMs show a default shape bias (47.9-73.8%) that exceeds their vision encoders and vision-only models but falls short of human levels (96%), with the LLM component rather than the encoder responsible for suppressing one visual cue. GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini , Qwen-VL , InternVL2 , LLaVA , LLaVA-NeXT / LLaVA 1.6 , MoE-LLaVA , InstructBLIP , Emu2 , CogAgent , CogVLM , UForm , CLIP / CLIP-ViT (LC) , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-416 Natural language prompts can steer the texture/shape bias in VLMs in both directions without significantly affecting accuracy, with texture-biased prompts more effective than shape-biased ones; this steering also generalizes to low/high-frequency bias. InternVL2 , LLaVA-NeXT / LLaVA 1.6 , Qwen-VL , Gemini IC-417 RLHF alignment reduces the creativity index of LLMs (GPT, Llama 2, OLMo) by an average of 30.1% at the verbatim level and 8.9% at the semantic level GPT-3 / GPT base , Llama 2 / Llama 2 base , OLMo / OLMo base IC-418 Matched n-grams in LLM outputs are concentrated in fewer reference documents than in human texts, indicating LLMs draw from a narrower set of sources GPT-3 / GPT base , Llama 2 / Llama 2 base , Tulu 2 , OLMo / OLMo base IC-419 CLIP ViT-B/16's layer-11 residual stream contains class-discriminative information in sparse SAE latent directions, and ablating class-specific top-k latents significantly degrades zero-shot classification accuracy CLIP / CLIP-ViT (LC) IC-420 CLIP ViT-B/16's SAE latent interpretability is depth-dependent: layer 11 encodes semantic object concepts while layers 2, 5, and 8 encode local shapes and attention patterns CLIP / CLIP-ViT (LC) IC-421 Sequential context-switching queries jailbreak Llama and Mistral models at 95% attack success rate Llama 3.1 , Llama-3.2-3B , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 2 / Llama 2 base , Vicuna , Cohere Command R IC-422 Safety fine-tuning in Llama models improves with parameter size but exhibits diminishing returns Llama-3.2-3B , Llama 3.1 IC-423 Llama-3.1-8B-Instruct achieves 97.8% accuracy as a zero-shot toxicity classifier on ToxiGen, outperforming Llama-3-Guard-1B and matching Llama-3-Guard-8B Llama 3.1 , Llama 3 IC-424 Structural in-context learning is transient in MultiBERTs and Pythia-1.4B, disappearing after early training MultiBERTs , Pythia IC-425 Pretrained GPT-2 Large fails at structural in-context learning on unseen tokens in a syllogism task GPT-2 IC-426 MultiBERTs exhibit a pushdown phenomenon where syntactic information migrates from later to earlier layers as training progresses MultiBERTs IC-427 Newer base models (post-November 2023) outperform older ones by 7.3 points on MMLU and 19.1 points on GSM8K controlling for pretraining compute, but this gap vanishes after fine-tuning all models on the same task-relevant data Pythia , Llama 2 / Llama 2 base , Llama 3 , Qwen1.5 , Gemma , OLMo / OLMo base , StableLM , Falcon , GPT-J , InternLM , OpenLLaMA , RedPajama , Baichuan , Skywork , Yi , ZiYA2 , MAP-NEO , Qwen 2 IC-428 Qwen 1.5 appears to Pareto-dominate Pythia and LLaMA 2 on MMLU and GSM8K, but after adjusting for test task training all three model families exhibit equivalent scaling Pythia , Llama 2 / Llama 2 base , Qwen1.5 IC-429 The point of emergence for MMLU shifts from approximately 1.3×10²² flops to 5.6×10²⁰ flops as models train on 64,000 task-relevant examples, and the log-linear fit R² improves from 0.632 to 0.950 Pythia IC-430 Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal GPT-2 , Llama 2 / Llama 2 base , Llama 3.1 IC-431 Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD dataset Llama 2 / Llama 2 base IC-432 56 LLMs from 19 families exhibit u-shaped scaling on hard questions and inverted-U scaling on easy questions, with the opposing trends explaining emergent ability stagnation Gemma , Llama 2 / Llama 2 base , RedPajama-INCITE , Yi , StableLM , MPT , Falcon , Pythia , Qwen , Qwen1.5 , BLOOM , DeepSeekMoE , OPT , GPT-Neo , CodeGen , XGLM , OpenLLaMA IC-433 LLMs exhibit a non-monotonic ID-OOD performance gap (generalization valley) that peaks at intermediate task complexity Qwen1.5 , Llama-3.2-3B , Llama 3 , Gemma 2 , Claude 3 , GPT-4o , O1 / OpenAI-o1-preview , Llama 3.1 , Qwen2.5 IC-434 The critical complexity at which LLMs over-rely on memorization shifts to higher task difficulty as model size increases Qwen1.5 , Llama 3.1 , Gemma 2 , Claude 3 , GPT-4o , O1 / OpenAI-o1-preview IC-435 Mistral-7B employs a less efficient algorithmic strategy (O(n²)) than Llama-3-8B (O([n², n³])) on probe tasks with multiple solution complexities Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 IC-436 API selection accuracy of 10 LLM-based agents degrades sharply as task complexity increases, with open-source models ≥70B matching closed-source on simpler tasks but lagging on the most complex Gemini 1.5 / Gemini Pro 1.5 , Llama 3 , Qwen 2 , Qwen2.5 , DeepSeek-2-Chat , DeepSeek-2-Coder , GPT-4o , GPT-3.5 / ChatGPT-3.5 , GLM-4 IC-437 Extracting parameters from user queries is harder for LLM-based agents than using outputs from previous actions, and less intelligent LLMs show steeper parameter-filling degradation with task difficulty Gemini 1.5 / Gemini Pro 1.5 , Llama 3 , Qwen 2 , Qwen2.5 , DeepSeek-2-Chat , DeepSeek-2-Coder , GPT-4o , GPT-3.5 / ChatGPT-3.5 , GLM-4 IC-438 All 10 LLM-based agents perform poorly at recognizing when they need to request input from the system or user, with overall accuracy between 30.55% and 55.18% Gemini 1.5 / Gemini Pro 1.5 , Llama 3 , Qwen 2 , Qwen2.5 , DeepSeek-2-Chat , DeepSeek-2-Coder , GPT-4o , GPT-3.5 / ChatGPT-3.5 , GLM-4 IC-439 Agent-specialized fine-tuned models (XLAM) significantly improve API selection over base models, but code-fine-tuned models (AgentLM) degrade performance, and no fine-tuning approach improves input recognition Agentlm , Xlam-R , Lemur-v1-70B / Lemur-70B-Chat-V1 , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-440 GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet show up to 25% skill-level accuracy gaps despite overall accuracies within 0.4% of each other GPT-4o , Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 IC-441 Skill-level improvements between model releases are highly uneven, with Claude 3.5 Sonnet gaining ~50% over Claude 3 Opus on law skills while Gemini improved most in math and science GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.5 / Gemini Pro 1.5 , Gemini 1.0 Pro , Claude 3.5 , Claude 3 IC-442 Routing each evaluation instance to the model strongest on its relevant skills yields a 3.2% accuracy gain over the best single model, with 3.5-6.8% gains on MMLU Pro GPT-4o , Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 IC-443 Model inconsistency on probing questions negatively correlates with skill-slice accuracy (r = -0.675), with models contradicting themselves more often on skills where they perform poorly GPT-4o , Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 IC-444 ViV1T produces non-differentiable population response representations when simulating mouse V1 experiments ViV1T IC-445 Stable Diffusion v1.5 generates nudity for 796 out of 4703 prompts in the I2P inappropriate prompts dataset Stable Diffusion IC-446 In Stable Diffusion v1.5, concept-generating neurons are localized in the second layer of FFNs, spanning less than 3% of FFN parameters, and are disentangled from object-generating neurons Stable Diffusion IC-447 CLIP ViT-B/16 produces noisy saliency maps and contains only 42 concept detectors, indicating poor visual interpretability CLIP / CLIP-ViT (LC) IC-448 CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasets CLIP / CLIP-ViT (LC) IC-449 CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255 CLIP / CLIP-ViT (LC) IC-450 LLaVA with the original CLIP encoder produces noisy, non-sparse attention maps that poorly localize to the objects described in generated text LLaVA IC-451 Transformer block coupling of Jacobian singular vectors positively correlates with benchmark performance across 30+ LLMs, more strongly than parameter count, depth, or embedding dimension Llama 3 , Llama 2 / Llama 2 base , Pythia , GPT-2 , Gemma , MPT , Falcon , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Phi-2 IC-452 Transformer block coupling is absent at initialization and increases persistently throughout training in Pythia 12B and 6.9B, with layer-wise locality emerging Pythia IC-453 Hidden representation trajectories in trained LLMs exhibit considerable linearity (mean LSS 4.25) compared to 6.54 at initialization, and linearity increases with training Llama 3 , Llama 2 / Llama 2 base , Pythia , GPT-2 , Gemma , MPT , Phi-2 IC-454 Most hidden trajectories in trained LLMs exhibit exponential growth in norm as a function of depth, a property that emerges with training Llama 3 , Llama 2 / Llama 2 base , Pythia , GPT-2 , Gemma , MPT , Phi-2 IC-455 Qwen-Audio 7B zero-shot underperforms a 128M parameter baseline on audio difference explanation across three evaluation scenarios Qwen-Audio IC-456 VLM decoders achieve near-random accuracy on VALSE image-sentence alignment while pairwise accuracy is much higher, indicating reliance on linguistic priors BakLLaVA , LLaVA-NeXT / LLaVA 1.6 , mPLUG-Owl3 IC-457 All four tested VLM decoders are heavily text-centric when generating answers, with text modality contributing 85-97% of the prediction signal BakLLaVA , LLaVA-NeXT / LLaVA 1.6 , mPLUG-Owl3 IC-458 Most VLM decoders show negative CC-SHAP on VALSE multiple-choice, indicating their explanations are less self-consistent than their answers, driven by a shift from text-dominant to image-dominant processing BakLLaVA , LLaVA-NeXT / LLaVA 1.6 , mPLUG-Owl3 IC-459 Qwen-2.5 models show that the privacy-utility tradeoff for differentially private steering improves with model size Qwen2.5 IC-460 Non-private activation steering of Llama-2-7B and Qwen-2.5-7B leaks membership information from the alignment dataset, while PSA reduces the empirical privacy loss Llama 2 / Llama 2 base , Qwen2.5 IC-461 Adding calibrated Gaussian noise to steering vectors (PSA) preserves alignment performance comparable to non-private mean steering across Llama-2-7B, Mistral-7B, Gemma-2-2B, and Qwen-2.5-7B Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma , Qwen2.5 IC-462 GPT-2 encodes toxicity in a low-dimensional linear subspace of its MLP layers, concentrated in higher layers GPT-2 IC-463 DPO's first-step gradients in GPT-2 are correlated with the toxic subspace, with stronger alignment in later layers and with more samples GPT-2 IC-464 Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate method Video-LLaVA , Llama-VID IC-465 GPT-4o achieves the highest weighted Cohen's kappa agreement with a reference-guided agent-debate method among all VLM judges, with scores exceeding 50 in several visual dimensions Video-LLaVA , Llama-VID , GPT-4o , InternVL2 IC-466 GPT-4o's evaluation reliability degrades when used as the final judge in a collective thought pipeline that aggregates reviews from less reliable VLMs Llama-VID , Video-ChatGPT , Video-LLaVA , GPT-4o IC-467 Llama-3.1-405B's standard speculative decoding verification rejects correct continuations from GPT-4o, Llama-3.1-8B, and human text, accepting only roughly two tokens before the first rejection for GPT-4o Llama 3.1 , GPT-4o IC-468 Llama-3.1-405B's last hidden layer embeddings of erroneous tokens contain a linearly detectable error signal that a simple logistic regression head can exploit to flag incorrect continuations Llama 3.1 IC-469 CLIP's global contrastive alignment causes attention on anatomically irrelevant regions in 3D CT, yielding limited zero-shot diagnostic accuracy (AUC 68.4 on 54 tasks) CLIP / CLIP-ViT (LC) IC-470 LOVT and MGCA, which use implicit cross-attention local alignment, show only marginal improvement over CLIP in 3D CT diagnosis (AUC 69.4 and 70.1 vs 68.4) LOVT , MGCA , CLIP / CLIP-ViT (LC) IC-471 Pythia models exceeding 100M parameters show a consistent leftward shift of the multifractal spectrum (increasing regularity) during training that is absent in the 14M and 31M variants Pythia IC-472 The degree of emergence metric derived from Pythia's internal structure positively correlates with benchmark performance across training epochs Pythia IC-473 ResNet-18 lacks a clear multifractal structure while ResNet-152 shows one with irregular shifts, and a 160M diffusion model exhibits lower degree of emergence than Pythia 160M ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , Stable Diffusion , Pythia IC-474 GPT-4, GPT-4o, and Llama-3.1-405B fail at knowledge classification and comparison without chain-of-thought GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Llama 3.1 IC-475 GPT-4, GPT-3.5, GPT-4o, and Llama-3.1-405B fail at inverse knowledge search regardless of prompting GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , GPT-4o , Llama 3.1 IC-476 GPT-4 and GPT-3.5 show strong positional bias in Chinese idiom character completion GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-477 Released LLMs achieve limited success rates as web agents on WebArena-Lite, with open-source models substantially below proprietary ones GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Llama 3.1 , GLM-4 , AutoWebGLM IC-478 GPT-4 and GPT-4V achieve approximately 71-73% accuracy in judging whether a web agent trajectory successfully completes a task GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-479 SAM's mask decoder exhibits attention drift to background or specific object parts under imprecise prompts, causing severe segmentation degradation SAM , SAM 2 IC-480 Pre-trained M-LLMs (GPT-4o, LLaVA-v1.6-34B, InternVL2-26B, Qwen2-VL-7B) produce imprecise tampering explanations when artifacts require fine-grained pixel-level analysis such as lighting or perspective inconsistencies GPT-4o , LLaVA-NeXT / LLaVA 1.6 , InternVL2 , Qwen2-VL IC-481 Llama-3.1-8B and four other released LLMs reorganize their internal representations to reflect in-context graph structure in a sudden two-phase transition as context length increases Llama 3.1 , Llama-3.2-3B , Gemma 2 IC-482 In Llama-3.1-8B, in-context graph structure is absent in early layers dominated by semantic priors and emerges clearly in deeper layers Llama 3.1 IC-483 When in-context graph structure conflicts with pretrained semantic priors, Llama-3.1-8B encodes the in-context structure in higher principal components while the semantic prior dominates the first two Llama 3.1 IC-484 Re-scaling the first 2-3 principal components of Llama-3.1-8B token representations causally shifts next-token predictions toward the target graph position Llama 3.1 IC-485 LLMs show a significant performance gap between Wikipedia-based factual multi-hop QA and counterfactual multi-hop QA, indicating reliance on memorized knowledge rather than reasoning from context GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Gemini , GPT-3 / GPT base , O1 / OpenAI-o1-preview , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Qwen 2 IC-486 LLMs achieve correct final answers through incorrect reasoning chains, inflating their apparent multi-step reasoning performance GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Gemini , GPT-3 / GPT base , O1 / OpenAI-o1-preview IC-487 Including sub-questions in the prompt improves LLM performance on multi-hop QA tasks GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Gemini , GPT-3 / GPT base , O1 / OpenAI-o1-preview IC-488 LLM performance degrades progressively as the number of reasoning hops increases, with error propagation from earlier sub-questions GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Gemini , GPT-3 / GPT base IC-489 State-of-the-art MLLMs fail at multi-step visual analogical reasoning, with best accuracy at 13% (Llama 3.2) on VOILA-WD and 29% (GPT-4o) on VOILA-ND, far below human performance of 71% and 70% GPT-4o , Llama 3.2 , Qwen2-VL , CogVLM2 , Seed-LLaMA-8B , MolmoE-7B , Emu2 IC-490 GPT-4o can identify visual relationships at 97% accuracy when given ground-truth descriptions but drops to 17% when asked to apply known relationships to new visuals, revealing a specific bottleneck in relational transfer GPT-4o IC-491 Presenting three images as a single collage rather than sequentially reduces MLLM accuracy by approximately 40% on the relationship application step GPT-4o , Qwen2-VL , LLaVA-OneVision IC-492 Least-to-most prompting consistently improves MLLM accuracy on the relationship application step compared to direct answering, with GPT-4o improving from 0.9% to 6.44% on VOILA-WD GPT-4o , CogVLM2 , Seed-LLaMA-8B IC-493 A linear direction in the input embedding space of Llama-2-7B-Chat, Llama-2-13B-Chat, Mistral-7B-Instruct-v0.3, and Phi-3-mini-128k predicts instruction-following success, generalizes across tasks but not instruction types, and can be used to improve adherence via representation engineering Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Phi-3 IC-494 The instruction-following dimension in Llama-2-7B-Chat and Llama-2-13B-Chat is more closely aligned with prompt phrasing than with task familiarity or instruction difficulty Llama 2 / Llama 2 base IC-495 All evaluated multimodal foundation models achieve average non-hallucination accuracy below 50% across six hallucination scenarios FLUX / FLUX1 , DALL·E 3 , DALL·E 2 , Stable Diffusion , Nova Pro , GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama-3-2-Vision , LLaVA-NeXT / LLaVA 1.6 , Gemini 1.5 / Gemini Pro 1.5 IC-496 GPT-4o achieves the highest location inference accuracy among evaluated models, reaching 98.16% for country, 60.23% for city, and 27.13% for zip code from street view images GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama-3-2-Vision , Nova Lite , Gemini 1.5 / Gemini Pro 1.5 IC-497 Text-to-image models experience performance drops exceeding 10% under adversarial prompts, with spatial reasoning being the most vulnerable task across all models Nova Canvas , FLUX / FLUX1 , DALL·E 3 , DALL·E 2 , GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Stable Diffusion , LLaVA-NeXT / LLaVA 1.6 IC-498 Multimodal foundation models exhibit severe group unfairness, with race and age biases more pronounced than gender bias in text-to-image models while gender bias is stronger in image-to-text models FLUX / FLUX1 , DALL·E 3 , DALL·E 2 , Stable Diffusion , GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini 1.5 / Gemini Pro 1.5 , Llama-3-2-Vision , Nova Canvas IC-499 ViT patch embeddings contain local semantic information beyond the [cls] token, as shown by performance degradation when restricting the output head to [cls] only or removing positional embeddings CLIP / CLIP-ViT (LC) , ViT , MAE , DINOv2 , SigLIP IC-500 GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B solve multi-step retrieval problems without fine-tuning, achieving near-perfect accuracy for chains of up to 5 steps GPT-4o , Claude 3.5 , Gemini 1.5 / Gemini Pro 1.5 , Llama 3.1 IC-501 Linear probes on middle-layer attention heads of Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 predict US lawmakers' DW-Nominate ideology scores with Spearman correlations around 0.85 Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Vicuna IC-502 Linear probes trained on US lawmaker ideology generalize to predict Ad Fontes media slant scores when the same models simulate news outlets Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Vicuna IC-503 Adding probe regression coefficients to attention head activations steers Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 toward more liberal or conservative generated text Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Vicuna IC-504 GPT-4o (gpt-4o-2024-08-06) rates the political slant of LLM-generated essays in close agreement with politically balanced human annotators GPT-4o IC-505 Adversarial attacks on Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT shift hidden representations along the negative refusal feature direction Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma IC-506 Restoring the refusal feature in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT causally disables all four tested adversarial attacks Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma IC-507 The refusal feature direction in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT ranks near the top among 100 perturbations for compromising model safety Llama 3 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Gemma IC-508 GPT-4 exhibits reduced preference consistency (0.66 vs 0.84) when the quality distinction between two responses is minimal GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-509 GPT-4 used as a preference labeler via prompt engineering yields alignment performance comparable to a task-specific 125M model GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-510 Qwen2-7B and Llama3-8B score near-random on textual temporal reasoning tasks while Qwen2-72B, Llama3-70B, and GPT-4o achieve near-perfect accuracy, showing temporal reasoning in LLMs is scale-dependent and emerges only above ~70B parameters Qwen 2 , Llama 3 , GPT-4o , LongVA-7B , ViLA-8B IC-511 LLaMA-2, Gemma, and Mistral all perform in-context density estimation via an adaptive kernel-like process, as revealed by their similar low-dimensional INPCA trajectories bounded between the geodesic and the Gaussian submanifold Llama 2 / Llama 2 base , Gemma , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-518 GPT-4o achieves 62.54% overall accuracy on MMWorld, the best among 15 MLLMs, while four open-source models perform below the 26.31% random-choice baseline GPT-4o , Claude 3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Gemini , Video-LLaVA , Video-Chat-7B , Chat-UniVi-7B , mPLUG-Owl , Video-ChatGPT , PandaGPT-7B , ImageBind-LLM-7B , X-InstructBLIP-7B , LWM-1M-JAX , Otter-7B , Video-LLaMA-2-13B IC-519 MLLMs exhibit different skill sets than humans, correctly answering expert-level questions that all three human annotators miss while failing on easy questions humans answer correctly GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Claude 3.5 , Gemini , Video-LLaVA , Video-Chat-7B , Video-ChatGPT , ImageBind-LLM-7B , PandaGPT-7B , Chat-UniVi-7B , Video-LLaMA-2-13B , X-InstructBLIP-7B , LWM-1M-JAX , Otter-7B , mPLUG-Owl IC-520 Temporal reasoning performance drops significantly across all 15 MLLMs when video frames are shuffled or reduced to one-fifth of the original count GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Claude 3.5 , Gemini , Video-LLaVA , Video-Chat-7B , Video-ChatGPT , ImageBind-LLM-7B , PandaGPT-7B , Chat-UniVi-7B , Video-LLaMA-2-13B , X-InstructBLIP-7B , LWM-1M-JAX , Otter-7B , mPLUG-Owl IC-521 MLLMs show asymmetric modality-specific perception, with Gemini Pro achieving 69.97% on visual-only questions but only 24.45% on audio-only, while Video-Chat outperforms ChatUniVi on audio despite worse visual scores Gemini , Video-Chat-7B , Chat-UniVi-7B , Video-LLaMA-2-13B , Otter-7B IC-522 LLMs perform correct example inference without inducing the correct rule, and this gap is robust to prompting methods, fact count, and scenario form GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , Claude 3.5 , Llama 3 , Llama 2 / Llama 2 base IC-523 LLMs rely on observed facts close to the test case in input feature space (neighbor-based reasoning) rather than on an abstract rule, and this effect is localized GPT-4o , Claude 3.5 , Llama 3 , Llama 2 / Llama 2 base IC-525 GPT-2 small's residual stream at layer 8 decomposes into two sub-spaces of approximately 25% and 75% of the dimensionality GPT-2 IC-526 GPT-2 small's first token position has residual stream norms more than an order of magnitude larger than all other positions GPT-2 IC-527 A 16 million latent sparse autoencoder substituted into GPT-4 yields a language modeling loss corresponding to 10% of GPT-4's pretraining compute GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-528 The knowledge localization assumption fails for a large fraction of facts in GPT-2, Llama2-7B, and Llama3-8B, with 77% of facts classified as inconsistent knowledge in Llama3-8B GPT-2 , Llama 2 / Llama 2 base , Llama 3 IC-529 For inconsistent knowledge in GPT-2, Llama2-7B, and Llama3-8B, the knowledge neurons are associated with the specific query rather than the fact, as shown by differential effects of suppressing or enhancing query-specific versus neighbor neurons GPT-2 , Llama 2 / Llama 2 base , Llama 3 IC-530 The attention module in GPT-2, Llama2-7B, and Llama3-8B plays a selective role in knowledge expression by activating specific knowledge neurons for a given query, as demonstrated by suppressing or enhancing attention scores at knowledge synapse positions GPT-2 , Llama 2 / Llama 2 base , Llama 3 IC-531 CONCH's zero-shot encoders cannot discriminate survival risk, achieving near-random concordance index on pathology whole-slide images CONCH IC-532 PLIP's zero-shot encoders produce random-guessing-level survival predictions and consistently underperform CONCH on pathology survival analysis PLIP , CONCH IC-537 Moirai's architectural enhancements (any-variate attention, multi-scale patch embedding, diverse mixture distribution) improve in-distribution forecasting but reduce out-of-distribution scalability relative to a simpler encoder-only baseline Moirai IC-538 Chronos-T5's discrete probability prediction approach yields very small power-law exponents on NLL, limiting its scalability, and its in-distribution gains do not extend to out-of-distribution data Chronos IC-539 GPT-4o in text-code-image mode achieves the highest scores on SCIMAGE but remains below 4 on all three evaluation dimensions, and all models degrade substantially on prompts requiring combined understanding types GPT-4o , Llama 3.1 , AutoTikZ / DataTikZ , DALL-E , Stable Diffusion IC-540 Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image models GPT-4o , Llama 3.1 , AutoTikZ / DataTikZ , Stable Diffusion , DALL-E IC-541 Code-based output (Python or TikZ) produces images with notably higher scientific style scores than direct image generation from Stable Diffusion and DALL-E GPT-4o , Llama 3.1 , Stable Diffusion , DALL-E IC-542 The PyTorch pretrained ResNet50 on ImageNet is vulnerable to (1, y)-ACE calibration attacks that increase ECE from 3.70% to 47.23% while preserving accuracy ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-547 Llama3.2 3B and Llama3.1 8B exhibit saturation events in which the top prediction, once it appears at a given layer, remains unchanged through all subsequent layers Llama-3.2-3B , Llama 3.1 IC-548 CLIP-B/32 exhibits progressively increasing layer-wise representation similarity in both its vision encoder and text encoder, and the pattern also holds across modalities CLIP / CLIP-ViT (LC) IC-549 All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBench GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , O1 / OpenAI-o1-preview , Claude 3.5 , Llama 3.1 , Llama 2 / Llama 2 base , Vicuna , Wizardlm , Qwen 2 , Qwen1.5 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Phi-3 , GLM-4 , InternLM-2.5-7B IC-550 Workflow generation performance scales with model size within families, but recently released 7B models outperform older 13B models Qwen 2 , Llama 3.1 , Llama 2 / Llama 2 base , Wizardlm , Vicuna , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , InternLM-2.5-7B IC-551 GPT-4's workflow generation performance declines as the number of nodes and edges in the workflow increases GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-552 GPT-4, Llama-3.1-8B, and Qwen-2-72B all improve on ALFWorld and WebShop when given a generated workflow as structured prior knowledge GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 3.1 , Qwen 2 IC-553 OpenCLIP ViT-B/16's image-text alignment score is a strong predictor of domain generalization accuracy, while perceptual similarity to LAION-400M pre-training data is a weaker predictor OpenCLIP IC-554 In Stable Diffusion 1.4, 1.5, 2.0, and 3.0, parameters with the smallest absolute values (below ~10^-3) do not contribute to the generative process, and this ineffectiveness is caused by stochastic training dynamics rather than architectural redundancy. Stable Diffusion IC-555 Large LLMs (GPT-3.5-turbo, Gemini 1.5 Flash, Llama3-70B, Mixtral 46.7B) exhibit reasoning errors and significant accuracy degradation on large-scale logical commonsense reasoning tasks with 32k+ rules, even when the knowledge base is complete and retrieval is ideal GPT-3.5 / ChatGPT-3.5 , Gemini 1.5 / Gemini Pro 1.5 , Llama 3 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , VERA IC-556 All 9 LLM-based guard models exhibit significant miscalibration with average ECE exceeding 10% across 12 public benchmarks for both prompt and response classification Llama Guard , Llama Guard 2 , Llama-Guard 3 , Aegis-Guard-Defensive , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , WildGuard IC-557 Guard models show significantly degraded calibration under jailbreak attacks, with prompt classification ECE substantially higher than response classification ECE Llama Guard , Llama Guard 2 , Llama-Guard 3 , Aegis-Guard-Defensive , WildGuard , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-558 Guard models exhibit inconsistent calibration when classifying responses from different response model types, with ECE varying by up to 39 percentage points within a single model Llama Guard , Llama Guard 2 , Llama-Guard 3 , Aegis-Guard-Defensive , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , WildGuard IC-559 Contextual calibration is most effective for prompt classification while temperature scaling is more effective for response classification, but no single post-hoc method fully resolves miscalibration Llama Guard , Llama Guard 2 , Llama-Guard 3 , Aegis-Guard-Defensive , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , WildGuard IC-560 Llama3-8B-Instruct reliably distinguishes its own outputs from human outputs in self-recognition tasks, while Llama3-8B base performs at chance, indicating the ability is acquired during post-training. Llama 3 IC-561 A linear direction in the residual stream at layer 16 of Llama3-8B-Instruct is causally necessary and sufficient for self-authorship claims: steering with it achieves 100% control over authorship assertions, and projecting it out reduces claims by 50-60%. Llama 3 IC-562 Applying the layer-16 self-recognition vector to input tokens (not output) of Llama3-8B-Instruct alters the model's perception of authorship, making it believe or disbelieve it wrote arbitrary texts in both individual and paired paradigms. Llama 3 IC-563 The self-recognition vector's activation in Llama3-8B-Instruct is organized across depth: early layers (4-6) show diffuse perceptual activation to self-written text (present in both chat and base models), while layers 14-16 show a sharp decision-related peak at the output token that is present only in the chat model with role tags. Llama 3 IC-564 The implicit attention matrices of Mamba, RWKV, and Griffin exhibit depth-dependent structure, with dependencies between distant tokens becoming more apparent in deeper layers Mamba , RWKV , Griffin IC-565 Phi-3's residual stream encodes format instructions as linear directions, evidenced by cosine similarity and vocabulary-space projections Phi-3 IC-566 Adding instruction-specific steering vectors to the residual stream improves instruction-following accuracy for Phi-3, Gemma 2 2B IT, Mistral 7B IT, and Gemma 2 9B IT across format, length, and word-specific constraints Phi-3 , Gemma 2 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-567 Steering vectors computed on instruction-tuned Gemma 2 models transfer to base Gemma 2 models, with cross-model steering outperforming same-model steering for Gemma 2 2B Gemma 2 IC-568 In Phi-3, word-exclusion steering vectors computed via difference-in-means project onto the vocabulary space with high logits for the excluded word, making them counterproductive Phi-3 IC-569 Mistral-7B-instruct-v0.1 achieves only F1 of 0.419 on zero-shot stance detection for the X-Stance German dataset, substantially below the fine-tuned BERT baseline (F1 0.693) Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-570 Gradient-based image jailbreaks optimized against single or ensemble VLMs are universal for the attacked model(s) but do not transfer to other VLMs, except between highly similar models Prismatic , Qwen-VL , DeepSeek-VL IC-571 Open-source VLMs (LLaVA, MiniGPT-4, InstructBLIP) are substantially more vulnerable to multimodal jailbreak attacks than Gemini-1.5-flash, with BAP attack ASR of 58–62% versus 40–41% LLaVA , MiniGPT-4 , InstructBLIP , Gemini IC-572 Bijection learning achieves state-of-the-art jailbreak ASR on frontier models, with peak ASR increasing with model capability Claude 3 , Claude 3.5 , GPT-4o IC-573 Model capabilities on MMLU degrade monotonically as bijection encoding complexity increases Claude 3 , Claude 3.5 , GPT-4o IC-574 Guard models fail to effectively mitigate bijection attacks even at capability parity with the target model Claude 3.5 , GPT-4o , Claude 3 IC-575 Four released LLMs (LLaMA-3.1-8B, Mistral-7B, Qwen2-7B, Yi-1.5-9B) can perform in-context learning on continuous vector representations projected into their embedding space, matching or outperforming few-shot ICL across text, time-series, graph, and fMRI tasks Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Qwen 2 , Yi IC-576 For 10-digit numerical function regression, vector-ICL consistently outperforms few-shot ICL with raw number inputs across all four LLMs because continuous representations avoid multi-token splitting Llama 3.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Qwen 2 , Yi IC-577 An encoder's text reconstruction performance positively correlates with its effectiveness in downstream vector-ICL classification tasks across 15 encoder-LLM-dataset configurations NV-Embed-v1 , SFR-Embedding-2-R , Stella-en-1.5b-v5 , GTR-T5-base IC-578 LLMs encode input text as linearly separable representations in forerunner token hidden states, emerging in early layers and enhanced by in-context demonstrations Llama 3 , Falcon IC-579 ICL hidden states exhibit positional bias: representations of the same input are more similar when the input appears at similar positions in the sequence Llama 3 IC-580 The 3-step ICL inference circuit (text encoding, semantics merge, feature retrieval) is a dominant causal mechanism, as ablating the corresponding attention connections significantly degrades ICL accuracy Llama 3 , Falcon IC-581 Induction heads for ICL operate on task-specific attention subspaces, with partial overlap across tasks, and the geometry of these subspaces explains demonstration saturation Llama 3 IC-582 Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIs InstructBLIP , mPLUG-Owl , Idefics , ViT , CLIP / CLIP-ViT (LC) , BLIP-2 , Llama 2 / Llama 2 base IC-583 Brain alignment in InstructBLIP and Idefics is organized by depth: middle layers align with higher visual regions while later layers align with early visual regions, whereas mPLUG-Owl shows later layers aligning with both InstructBLIP , mPLUG-Owl , Idefics IC-584 Most brain-explained variance is shared across task instructions, with image captioning (IC) acting as an umbrella category showing high overlap with VQ and CR but lower overlap with iu2 and sr InstructBLIP IC-585 MLLMs effectively capture count-related and recognition-related visual concepts with distinct brain alignment patterns, but produce similar alignment patterns for color, positional understanding, and general scene understanding InstructBLIP , mPLUG-Owl , Idefics IC-586 Symbolic distance (number of reasoning steps) is the primary bottleneck for relational reasoning in LLMs, not total context length Gemma , Mixtral , Gemini , GPT-4o IC-587 Real-world knowledge acts as a shortcut in LLM relational reasoning, causing worse-than-chance performance on logically valid but factually incongruent statements Gemma , Mixtral , Gemini , GPT-4o IC-588 Topologically ordered context improves relational reasoning over random ordering across nearly all LLMs Gemma , Mixtral , Gemini , GPT-4o IC-589 Flavor text (non-essential descriptive language) degrades relational reasoning in most LLMs, but GPT-4o is robust to it Gemma , Mixtral , Gemini , GPT-4o IC-590 Tuning only the identified safety neurons (SN-Tune) reduces harmful scores of instruction-tuned and base models by over 90 points while preserving general capability. Vicuna , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 2 / Llama 2 base , Llama 3 IC-591 Downstream fine-tuning on GSM8K degrades safety of Llama2-7b-chat and Mistral-7b-instruct-v0.2, but RSN-Tune partially preserves safety by protecting non-overlapping safety neurons. Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-592 The log-likelihood layer in LLaMA-2-7B, LLaMA-2-7B-Chat, Vicuna-7B, and Mistral-7B-Instruct produces factually incorrect answers on TruthfulQA MC1 (817 samples) due to a misalignment between the output distribution and internal attention head representations, with LM-to-head-norm accuracy gaps of 24.23 to 40.68 points. Llama 2 / Llama 2 base , Vicuna , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Zephyr-7B-beta , Llama 3 , Gemma 2 IC-593 The L2 norms of attention heads in Mistral-7B-Instruct and LLaMA-2-7B correlate with truthfulness, spiking by up to 83% at token positions of factual proposition completions and pertinent factual associations, and this correlation is specific to multi-headed attention representations rather than query, key, value, output, or FFN norms. Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 2 / Llama 2 base IC-594 In LLaMA-2-7B, the truth-correlated attention heads are concentrated after layer 9, with two functional types (structural and associative) evenly distributed throughout the upper portions of the model, showing no further depth-dependent specialisation within that region. Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-595 Gemini-pro has knowledge gaps on specific topics (Permian extinction, Fordism) causing it to perform far below its average rank on existing benchmarks Gemini , Claude 2.0 , GPT-3.5 / ChatGPT-3.5 IC-596 Multiple released LLMs fail to refuse harmful prompts disguised as historical or philosophical discussions, with GPT-4o and Mixtral showing the lowest refusal rates GPT-4o , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3 , Claude 3 IC-597 Most LMMs exhibit systematic class bias in synthetic data detection, with GPT-4o biased toward classifying text as real and 3D as AI-generated GPT-4o , Claude 3.5 , Gemini 1.5 / Gemini Pro 1.5 , InternVL2 , Qwen2-VL , LLaVA-OneVision IC-598 GPT-4o's synthetic image detection accuracy drops sharply on specialized domains (satellite 45.0%, medical 54.3%) compared to common image types (object 84.4%, person 84.4%) GPT-4o , Qwen2-VL IC-599 All evaluated audio LMMs perform at or near random chance (44.4%–51.2%) on synthetic audio detection, while humans achieve 69.2% Qwen-Audio , SALMONN , OneLLM , Gemini 1.5 / Gemini Pro 1.5 , AASIST IC-600 Chain-of-thought prompting improves most LMMs on synthetic detection but degrades LLaVA-ov-7b from 56.6% to 18.8%, while GPT-4o performs well without it (64.1% baseline) GPT-4o , LLaVA-OneVision , InternVL2 , Qwen2-VL , Gemini 1.5 / Gemini Pro 1.5 , Claude 3.5 IC-601 Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreases Qwen 2 , Llama 3 IC-602 Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue grows Qwen 2 , Llama 3 IC-603 SIREN's embedding layer is not effectively optimized by gradient descent, so its embedding frequencies must be set as a hyperparameter SIREN IC-604 CLIP ViT-B/32's CIFAR-10 image embeddings approximately satisfy a multi-cluster structure with near-orthogonal class-mean features CLIP / CLIP-ViT (LC) IC-605 Six SOTA LLMs are vulnerable to composable jailbreak attacks, with maximum attack success rates ranging from 44% to 94% GPT-3.5 / ChatGPT-3.5 , GPT-4o , Claude 3 , Llama 3 IC-606 The relationship between model size and jailbreak vulnerability is reversed between Anthropic and Meta model families Claude 3 , Llama 3 , GPT-3.5 / ChatGPT-3.5 , GPT-4o IC-607 CLMBR-T-BASE's clinical prediction performance degrades as patient EHRs become more repetitive or irregular CLMBR-T-BASE IC-608 Token trajectories in GPT-2, Llama 2 7B, Mistral 7B, and Llama 3.2 models cluster on a low-dimensional manifold and follow a linear drift plus Gaussian noise dynamics GPT-2 , Llama 2 / Llama 2 base , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3.2 IC-609 The last transformer layer of Mistral 7B, Llama 3.2 1B, and Llama 3.2 3B shows anomalous trajectory statistics inconsistent with the linear drift-plus-noise pattern of intermediate layers Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Llama 3.2 IC-610 In LLaMA2-7B-Chat, RAG hallucinations are causally driven by copying heads losing external context information during generation and by knowledge FFNs in mid-to-upper layers over-adding parametric knowledge to the residual stream Llama 2 / Llama 2 base , Llama 3 IC-620 CAIT-S/24 (a ViT) attends more to high-frequency image features than ResNet-101 (a CNN), as revealed by frequency-domain attribution visualization CAIT-S/24 , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-626 Llama 7B retains 95% of its common sense reasoning performance when compressed to 2 GB using JLCM LLaMA IC-627 ResNet 18 and ViT B/16 retain significant ImageNet accuracy at 2–3 bit weight compression via JLCM ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , ViT IC-632 BLIP-2 succeeds on only 5 out of 100 advanced compositional vision-language tasks BLIP-2 IC-633 BLIP-2, LLaVA, and mPLUG-Owl show a trade-off between caption length and hallucination rate on COCO BLIP-2 , LLaVA , mPLUG-Owl IC-634 Video-ChatGPT achieves the highest correctness, detail, and contextual scores among four video understanding baselines Video-ChatGPT , VideoChat , LLaMA-Adapter , Video-LLaMA IC-635 InstructBLIP achieves the highest overall MMBench score (44.0) among five evaluated vision-language models InstructBLIP , LLaVA , VisualGLM , OpenFlamingo IC-636 Syntactic phenomena (determiner-noun and subject-verb agreement) localize to the same topmost-layer MLP neurons as factual information in BERT, GPT-2, and Llama-2 BERT , GPT-2 , Llama 2 / Llama 2 base IC-637 KN edit (neuron suppression) has low reliability, overturning at most 5.2% of BLIMP categorical predictions and achieving only 1.66%–47.86% reliability on factual tasks BERT , T5 , GPT-J IC-638 ROME editing on GPT-2 XL and Llama-2 7B achieves high reliability but fails under bijective symmetry (23.71%–33.64%) and synonymous invariance (52.35%–58.36%) criteria GPT-2 , Llama 2 / Llama 2 base IC-639 The causal tracing pattern of MLP at early layers and attention at late layers is not stable across factual and syntactic phenomena in GPT-2 XL GPT-2 IC-642 CLIP can infer contextual attributes (orientation, illumination, etc.) from images with approximately 74% accuracy on a binary task CLIP / CLIP-ViT (LC) IC-643 Conditioning CLIP on correct contextual attributes in the text prompt improves zero-shot classification accuracy across 13 image transformations CLIP / CLIP-ViT (LC) IC-644 CLIP relies on spurious features (background) as a shortcut in zero-shot classification, and conditioning on the correct background reduces this reliance CLIP / CLIP-ViT (LC) IC-645 CLIP, PickScore, and HPSv2 text embeddings share a common direction (cone effect) that captures text-irrelevant preferences, and the orthogonal component c⊥p better measures T2I alignment; CLIP's untrained common direction makes it ineffective for reward fine-tuning CLIP / CLIP-ViT (LC) , PickScore , HPSv2 IC-646 DINOv2, DeiT-III, and OpenCLIP repurpose approximately 2% of patch tokens in low-informative background areas as internal registers, discarding local patch information while aggregating global image information; DINO does not exhibit this behaviour DINOv2 , DINO , DeiT-III , OpenCLIP , MAE IC-647 DINOv2's feature-map artifacts cause it to be incompatible with the LOSt unsupervised object discovery method, scoring far below DINO DINOv2 , DINO , DeiT-III , OpenCLIP IC-662 Translation ability in BLOOM models surges at approximately one-sixth of pre-training and then plateaus, with consistent dynamics across model sizes from 560M to 7.1B BLOOM IC-663 OPT-1.3B, LLaMA-7B, and Aquila-7B encode sparse Harsanyi interactions, with only 29-51 salient interactions out of 1024 possible on SQuAD sentences OPT , LLaMA , Aquila-7B IC-666 DINOv2, CLIP-vision, and VGG-19 representations all align with MEG brain responses, with DINOv2 showing particularly high retrieval performance for late brain activity after image offset DINOv2 , CLIP / CLIP-ViT (LC) , VGG / VGG13 IC-667 CLIP's intermediate layer features encode object boundaries recoverable by k-means clustering, a property absent in shallow and deep layers CLIP / CLIP-ViT (LC) IC-668 SAM's edge-oriented segmentation yields high recall but very low precision because it cannot distinguish object boundaries from interior edges SAM IC-669 DINOv2's features, when clustered, produce smooth semantic regions but lack instance-level boundary delineation DINOv2 IC-673 Trained depthwise convolutional kernels in DS-CNN architectures converge to identifiable DoG-like patterns, with over 95% of ConvNeXtV2 and over 90% of ConvNeXt filters classifiable into a small set of clusters ConvNeXtV2 , ConvNeXt , MoGANet , ConvMixer , EfficientNet , MobileNetV3 , MNASNet , ReplkNet-XL IC-677 CLIP ViT's image representation is primarily constructed by the last 4 MSA layers, with MLPs and early MSA layers contributing negligibly OpenCLIP IC-678 Specific attention heads in CLIP ViT-L's last 4 layers encode specific image properties (color, shape, location, counting, texture) that are linearly recoverable via text directions OpenCLIP IC-679 CLIP relies on background/location as a spurious cue for bird classification, and ablating geolocation heads improves worst-group accuracy by 25.2% OpenCLIP IC-680 CLIP ViT's image token contributions are spatially localized to match described content, enabling zero-shot segmentation that outperforms existing CLIP-based methods OpenCLIP IC-681 GPT-3.5 exhibits positional bias when judging which of two LLM responses is superior GPT-3.5 / ChatGPT-3.5 IC-682 GPT-3.5 and GPT-4 achieve F1 scores of 0.5820 and 0.6180 respectively on pairwise response quality evaluation against human annotations GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-683 LLaMA-7B, LLaMA-30B, Vicuna-7B, and Vicuna-13B achieve low accuracy (12.11% to 42.24%) in zero-shot and few-shot log-likelihood response evaluation LLaMA , Vicuna IC-684 CLIP ViT-B/32 misclassifies 99% of forest satellite images as ocean when the word 'ocean' is overlaid as text CLIP / CLIP-ViT (LC) IC-685 CLIP ViT-B/32 with a linear probe relies on gender as a spurious correlation for hair color, achieving only 15.85% accuracy on female gray hair CLIP / CLIP-ViT (LC) IC-686 Adversarial perturbations alter CLIP ViT-B/32's token representations most strongly starting around layer 10 CLIP / CLIP-ViT (LC) IC-687 GPT-3.5+ models exhibit a gambler's fallacy bias and generate low-complexity sequences when asked to produce random binary sequences GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-688 GPT-3.5-turbo-instruct-0914 shows sharp phase transitions in in-context learning of simple formal languages, transitioning from random generation to deterministic pattern repetition as context length increases GPT-3.5 / ChatGPT-3.5 , GPT-3 / GPT base IC-689 Subjective randomness generation and sharp ICL transitions emerge only in larger or reward-fine-tuned models, absent in earlier GPT-3 variants and smaller open-source models GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , text-ada-001 , Llama 2 / Llama 2 base , Tulu 2 , Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b IC-698 GPT-3.5-turbo's CoT reasoning errors are correlated across different demonstration sets, while PoT errors are less correlated GPT-3.5 / ChatGPT-3.5 IC-699 Llama2-13b produces significantly less consistent answers than GPT-3.5-turbo on complex reasoning tasks, making it unsuitable as a weaker LLM in a cascade Llama 2 / Llama 2 base , GPT-3.5 / ChatGPT-3.5 IC-700 GPT-4's reasoning accuracy degrades when provided with incorrect hints from a weaker model GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-709 Swin Transformer parameter redundancy is depth-dependent, with higher layers retaining fewer parameters than lower layers under module-aware pruning Swin Transformer IC-713 Llama-2-7B-Instruct exhibits a reproducible failure mode in book-length summarization: high repetition and complete inability to perform incremental updating Llama 2 / Llama 2 base IC-714 For GPT-4 book-length summaries, human annotators prefer incremental summaries for detail (83% vs 11%) but hierarchical for structure (59% vs 35%), logic (53% vs 38%), and overall (54% vs 44%), showing coherence and human preference are not aligned GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-715 Factual information deleted from GPT-J, LLaMA-2, and GPT-2-XL via ROME or MEMIT remains linearly recoverable from intermediate hidden states, with up to 89% extraction success at budget b=20 GPT-J , Llama 2 / Llama 2 base , GPT-2 IC-716 Factual information deleted from GPT-J, LLaMA-2, and GPT-2-XL via ROME or MEMIT is recoverable by sampling outputs on automatically generated rephrased prompts, with up to 56% extraction success at budget b=20 GPT-J , Llama 2 / Llama 2 base , GPT-2 IC-717 Llama-2-Chat's evaluation capability does not improve monotonically with model size Llama 2 / Llama 2 base IC-718 GPT-4 achieves 0.882 Pearson correlation with human evaluators on 45 customized score rubrics while GPT-3.5-Turbo achieves only 0.392 GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-719 Llama-2-Chat achieves reasonable human-preference accuracy (51.78-53.67%) as a prompted reward model without specific reward-model training Llama 2 / Llama 2 base , StanfordNLP Reward Model , ALMOST Reward Model , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-721 The 72-head entity tracking circuit identified in Llama-7b achieves high faithfulness in Vicuna-7b and Goat-7b without any modification to the circuit graph LLaMA , Vicuna , GOAT-7B IC-722 Entity tracking in Llama-7b is implemented by detecting and transmitting the positional information of the correct entity, with distinct head groups for position detection, transmission, and value fetching LLaMA , Vicuna , GOAT-7B IC-723 The entity tracking performance gap between Goat-7b and Llama-7b is primarily attributable to enhanced positional information in the value fetcher and position transmitter heads LLaMA , GOAT-7B IC-730 CLIPCap and BLIP-2 produce degraded alt-text on Twitter social media images, with BLEU@4 of 0.372 and 0.111 respectively CLIPCap , BLIP-2 IC-731 CLIP (ViT-B/32) achieves only 17.5 recall on video-text temporal alignment because it was trained on images and lacks video dynamics CLIP / CLIP-ViT (LC) IC-732 Off-the-shelf PyTorch ResNet classifiers are better calibrated than fine-tuned U-Net classifiers at high noise levels in the diffusion reverse process ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-733 CoT prompting improves factual accuracy for instruction-tuned LLMs but degrades it for non-instruction-tuned LLMs such as OPT, BLOOM, and LLaMA OPT , BLOOM , LLaMA , Vicuna , ChatGLM-6B / ChatGLM-6b-2 , FLAN-T5 , GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 IC-734 GPT-3.5-turbo's factual verification F1 decreases as the number of reasoning hops required to validate a claim increases GPT-3.5 / ChatGPT-3.5 IC-735 GPT-3.5-turbo's factual verification performance drops substantially under adversarial modifications, with man-made adversarial examples causing the largest decline GPT-3.5 / ChatGPT-3.5 IC-736 Vicuna-13B outperforms Vicuna-7B on factual knowledge tasks by 5.4% on average Vicuna IC-737 CLIP ViT-B/16 binarized dot products yield 0.50–0.58 accuracy on binary concept presence queries across five image classification datasets CLIP / CLIP-ViT (LC) IC-738 BLIP-2 ViT-G FlanT5XL achieves 0.70–0.87 zero-shot accuracy on binary concept presence queries, competitive on most datasets but weaker on fine-grained CUB-200 BLIP-2 IC-739 GPT-3.5-turbo-0613 combined with CLIP produces more faithful concept-salience pseudo-labels than LLaMA-2-13B-Chat, InstructBLIP, or LLaVA-1.5B on most of five datasets GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 2 / Llama 2 base , InstructBLIP , LLaVA , CLIP / CLIP-ViT (LC) IC-742 ResNet-50-BN on Waterbirds relies on background as a spurious feature for classification, and this shortcut is invisible to entropy-based confidence metrics ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN IC-744 The decoded vocabulary of a function vector often reflects the task's output space, but reconstructing a vector that matches this vocabulary distribution does not recover the FV's full causal effect. GPT-J IC-745 Function vectors for simple list-oriented tasks can be algebraically combined via addition and subtraction to produce new vectors that trigger composed tasks, sometimes outperforming 10-shot ICL. GPT-J , Llama 2 / Llama 2 base IC-746 RoBERTa-Large pretrained with different mask ratios exhibits a sweet spot in downstream accuracy on QNLI and SST-2 RoBERTa / RoBERTa-L IC-747 SOTA pruning methods (SparseGPT, Wanda, magnitude) cause significant degradation on knowledge-intensive tasks for Vicuna and Llama models at 25-30%+ unstructured sparsity, and fail completely for n:m structured sparsity Vicuna , LLaMA , Llama 2 / Llama 2 base IC-748 Pruned LLMs at ≥50% sparsity remain robust in-context retrievers and summarizers, with Vicuna-7B matching up to ~40% sparsity and Vicuna-13B up to ~50% sparsity in open-book settings Vicuna IC-749 Compressed Vicuna-13B at 46.16% sparsity (matching 7B parameter count) achieves lower MMLU accuracy than dense Vicuna-7B, indicating large-sparse models do not outperform small-dense at matched size Vicuna IC-750 Open-source models without safety training are significantly more vulnerable to jailbreak attacks than safety-aligned proprietary models Vicuna , Alpaca , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Llama 2 / Llama 2 base IC-751 Arena-Hard-200 reveals larger performance gaps between open and proprietary LLMs than MT-Bench GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Claude Instant 1 , Vicuna , Llama 2 / Llama 2 base , Wizardlm IC-752 GPT-4's win rate over GPT-3.5-turbo is 52% on the top-50 most challenging prompts but only 22% on the bottom-50 easiest prompts GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-761 GPT-3.5-turbo and GPT-4 are susceptible to specific circulating jailbreaking prompts, with 'jailmommy' achieving a 71.16% success rate in producing toxic outputs GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-762 GPT-4 and GPT-3.5 outperform humans in generation but underperform in discriminative (selective) evaluation across 10 of 13 language tasks GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-763 CLIP and OpenCLIP fall short of human discriminative accuracy on vision tasks, with performance dropping substantially under hard negatives CLIP / CLIP-ViT (LC) , OpenCLIP IC-764 GPT-4 and GPT-3.5 make frequent errors answering questions about their own generated text, underperforming humans in interrogative evaluation GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-765 BLIP-2, BLIP, InstructBLIP, Bard, and BingChat fall short of human accuracy in answering questions about Midjourney-generated images BLIP-2 , BLIP , InstructBLIP , Bard , BingChat IC-766 Multiple Real-SR methods fail to outperform a small FSRCNN network on the majority of 100 representative degradation cases SRResNet , DASR , RD-SR , ESRGAN IC-767 BSRNet outperforms RealESRNet on the majority of degradation cases, reversing the ranking obtained from a single random test set BSRNet , RealESRNet IC-768 MMRealSR exhibits the most consistent performance across degradation cases while SwinIR achieves the highest performance on cases it handles well, among GAN-based Real-SR methods MMRealSR , SwinIR IC-783 Stable Diffusion v1.5's conditional probability pθ(x|c) is heavily biased by the unconditional probability pθ(x), making it unreliable as a condition-alignment metric Stable Diffusion IC-784 Pre-trained scoring models (CLIP Score, HPS, Image Reward, Pick Score) underperform on domain-specific fine-tuned diffusion models CLIP / CLIP-ViT (LC) , HPS , ImageReward , PickScore , Van Gogh Diffusion IC-785 Different released diffusion models produce images with distinguishable probability signatures, enabling source attribution Dreamlike Photoreal 2.0 , OpenJourney , Stable Diffusion IC-786 CLIP-ViT (LC) achieves 0.87 accuracy and 0.91 average precision on fake image detection CLIP / CLIP-ViT (LC) IC-788 CLIP ViT-B/32 fails to retrieve the correct image even when the generated target caption is well-aligned with the ground-truth image CLIP / CLIP-ViT (LC) IC-789 CLIP retrieval quality in zero-shot compositional image retrieval scales log-linearly with model size from approximately 150M to 2.5B parameters CLIP / CLIP-ViT (LC) , OpenCLIP IC-790 GPT-4 outperforms GPT-3.5-turbo, Vicuna-13B, and Llama2-70B for generating target captions in zero-shot compositional image retrieval GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Vicuna , Llama 2 / Llama 2 base IC-791 BLIP-2, BLIP, and COCA produce captions of comparable quality for zero-shot compositional image retrieval BLIP-2 , BLIP , COCA IC-795 1D subspaces of MLP activations found by DAS in GPT-2 Small (IOI) and GPT-2 XL (factual recall) produce apparent causal effects that are interpretability illusions driven by causally disconnected components activating dormant pathways GPT-2 IC-796 GPT-2 Small MLP weight matrices are full-rank across all 12 layers and residual stream features are linearly recoverable from post-GELU MLP hidden activations, providing the structural conditions for the subspace patching illusion GPT-2 IC-798 GPT-4 achieves 61.6% on MATH with tool-integrated reasoning prompting, exceeding PaL (51.8%) and CoT (42.5%) GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-799 WizardMath-70b scores lower than base Llama-2-70b on TabMWP (49.8% vs 57.5%), indicating degraded OOD generalization from rationale-based fine-tuning WizardMath , Llama 2 / Llama 2 base IC-807 Language-conditioned robot policies RT-1 and RT-2 fail to generalize to unseen manipulation tasks, achieving only 16.7% and 11.1% success rates respectively across 7 novel skills RT-1 , RT-2 IC-808 Sparse autoencoder features in Pythia-70m's residual stream are more interpretable than PCA, ICA, random, and default-basis directions, with the advantage declining from early to late layers Pythia IC-809 Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergence Pythia IC-810 Individual sparse autoencoder features in Pythia-70m-deduped are monosemantic and have predictable causal effects on output logits, as demonstrated by an apostrophe feature whose ablation primarily suppresses the 's' token Pythia IC-811 Alpaca's 52k instruction-tuning data is predominantly low-quality (only 17.75% score ≥ 4.5 on accuracy), yet the full 52k data still yields higher MMLU scores than the filtered 9k subset for both 7b and 13b variants Alpaca IC-812 InstructGPT (text-davinci-003) reduces content diversity in co-written essays while GPT-3 (davinci) does not, and the effect is attributable to the model's own less diverse text contributions InstructGPT , GPT-3 / GPT base IC-815 RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B models Pythia , LLaMA IC-816 RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B models Pythia , LLaMA IC-817 RLHF on general-purpose preference data improves machine ethics performance in Pythia and Llama-7B models Pythia , LLaMA IC-818 RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B models Pythia , LLaMA IC-824 GPT-3 models of all sizes (350M to 175B) can reverse name-description associations in-context with near-perfect accuracy, showing the reversal curse is a property of training rather than reasoning GPT-3 / GPT base IC-827 LLMs exhibit distinct psychological profiles that differ from human norms and vary by model size and version GPT-3 / GPT base , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 2 / Llama 2 base IC-828 Jailbreaking GPT-4 via cipherchat shifts its psychological profile toward human norms and reduces emotional intelligence scores GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-829 Role assignment to GPT-3.5-turbo produces role-consistent changes in psychological profiles and task performance, validating the psychometric scales GPT-3.5 / ChatGPT-3.5 IC-834 CLIP zero-shot predictions exhibit high equal opportunity difference when target and sensitive attributes are intrinsically dependent CLIP / CLIP-ViT (LC) IC-835 CLIP zero-shot predictions exhibit large worst-group accuracy gaps due to spurious correlations on Waterbirds and CelebA CLIP / CLIP-ViT (LC) IC-836 CLIP zero-shot predictions exhibit demographic bias on FairFace when using attribute-unrelated text prompts CLIP / CLIP-ViT (LC) IC-837 CLIP ViT-L/14 embeds more target-attribute information and less sensitive-attribute information than CLIP ResNet-50 on Waterbirds CLIP / CLIP-ViT (LC) IC-840 GPT-2 Small's name mover heads exhibit disrupted attention patterns under out-of-distribution Gaussian noise corruption GPT-2 IC-847 CLIP ViT-B/32 CLIPScore achieves only ρ=0.276 / τ=0.191 correlation with human 1-5 likert T2I alignment ratings on TIFA160 CLIP / CLIP-ViT (LC) IC-848 GPT-3.5 achieves 98.3% precision and 96.0% recall for automatic question-tuple matching but makes errors when questions differ in wording yet are semantically unique GPT-3.5 / ChatGPT-3.5 IC-849 GPT-2 and T5-base exhibit vanishing expected gradients under RFT for inputs with small reward standard deviation, prevalent in 3 of 7 GRUE datasets, causing RFT to underperform SFT GPT-2 , T5 IC-850 A partial SFT phase (40% of steps, 1% of samples) before RFT allows GPT-2 and T5-base to reach 96% of the reward achieved with full SFT+RFT, by reducing the number of inputs with vanishing gradients GPT-2 , T5 IC-852 Intrinsic self-correction without external feedback consistently degrades reasoning accuracy across GPT-3.5-turbo, GPT-4, GPT-4-turbo, and LLaMA-2-70B-chat GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Llama 2 / Llama 2 base IC-853 Multi-agent debate with GPT-3.5-turbo-0301 does not outperform self-consistency at equivalent inference cost on GSM8K GPT-3.5 / ChatGPT-3.5 IC-854 The apparent self-correction improvement in constrained generation (Madaan et al., 2023) is an artefact of a sub-optimal initial prompt, not a genuine model capability GPT-3.5 / ChatGPT-3.5 IC-858 The choice of graph encoding method significantly changes LLM accuracy on graph reasoning tasks, with incident encoding outperforming adjacency by up to 34 percentage points on connected nodes PaLM 62B , PaLM 2 , GPT-3.5 / ChatGPT-3.5 IC-859 LLMs rely on learned priors about graph properties (cycles exist, edges are absent) rather than analyzing the specific graph structure, causing below-majority-baseline performance and extreme structure-dependent accuracy PaLM 62B IC-860 Larger PaLM 2 models (xxs to l) show progressively better graph reasoning, but even the largest variant fails to beat the majority baseline on edge existence PaLM 2 IC-861 LLMs achieve near-zero accuracy on the disconnected nodes task, indicating an inability to reason about the absence of edges in a graph PaLM 62B IC-866 A prefix applied to Llama-7B's first attention layer preserves the relative attention distribution over content positions and only adds a constant-direction bias to the attention block output LLaMA IC-867 In GPT-2 prefix-tuned on the emotion dataset, attention over prefix positions is nearly constant across inputs, collapsing the effective bias subspace to a single direction in most layers GPT-2 IC-868 Skill-Mix performance degrades with increasing k, and within the Llama-2 family the saturation point increases with model size Llama 2 / Llama 2 base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Falcon , Xwin-LM-70B-v0.1 , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , Qwen , TigerBot-70B-Chat IC-869 Models ranking highly on popular LLM leaderboards perform worse than Llama-2-70b-chat on Skill-Mix, suggesting cramming for the leaderboard at the expense of general-purpose text generation Falcon , Xwin-LM-70B-v0.1 , Qwen , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , TigerBot-70B-Chat , Llama 2 / Llama 2 base IC-870 GPT-4's performance on Skill-Mix(k=5) and Skill-Mix(k=6) provides probabilistic evidence of generating novel skill-topic combinations not present in training data GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-871 Llama-2-70b-chat as a grader is more generous than GPT-4 and systematically gives higher scores to Llama-2 family outputs Llama 2 / Llama 2 base , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 IC-879 Code Llama outperforms Llama-2 on coding (HumanEval) and mathematical (GSM8K) reasoning at both 7B and 13B scales Llama 2 / Llama 2 base , Code Llama IC-887 HuggingGPT's in-context task-model assignment always selects the same model regardless of input question or task type HuggingGPT IC-888 SLIMG fails on heterophily graphs for link prediction because it cannot properly measure node similarity of heterophily embeddings SLIMG IC-894 FLAN-T5 base's average per-token probability increases with token index during generation on WMT translation tasks FLAN-T5 IC-895 Across all FLAN-T5 sizes, BLEURT scores show a negative correlation with prediction length on WMT translation tasks FLAN-T5 IC-897 GPT-4 underperforms codex on Spider text-to-SQL with few-shot prompting, attributed to its zero-shot tuning GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , code-davinci-002 IC-898 GPT-3.5-turbo and GPT-4 are overconfident in their initial code predictions when unit test execution is unavailable GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-899 GPT-4, when used as a blind pairwise evaluator, exhibits the same style-over-factuality preference as human crowdworkers GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-900 BLIP-2, MiniGPT-4, and LLaVA-1.5 show degraded zero-shot VQA accuracy on underspecified questions, with absolute improvements of 1.14–7.94% when questions are augmented with visually-grounded details BLIP-2 , MiniGPT-4 , LLaVA-1.5 / LLaVA-v1.5 IC-901 BLIP-2's LLM-only VQA performance improves with more specified questions while the image remains essential, revealing asymmetric strength between the LLM and vision components BLIP-2 IC-902 BLIP-2 and MiniGPT-4 confidence-based question selection underperforms the original question for paraphrased candidates but succeeds for semantically enriched REPARe questions BLIP-2 , MiniGPT-4 IC-903 Knowledge editing performance (ES, GS, LS) improves as model scale increases from GPT-2 (124M) to T5-XL (2.8B) to GPT-J (6B) across all editing methods GPT-2 , T5 , GPT-J IC-904 Stable Diffusion XL generates non-empty cups when prompted for 'empty cup' Stable Diffusion IC-913 In OPT-2.7B, Pythia-70M/1.4B/6.9B, and BERT-base, the stable rank of MLP lower layers shows a drop-and-bounce pattern during training that is more salient in top layers while bottom layers show suppressed dropping curves OPT , Pythia , BERT IC-914 In Pythia models (70M through 2.8B), BERT-base, OPT-6.7B, LLaMA-2-7B, and ViT-Huge, the MLP out-projection vectors are almost orthogonal throughout training Pythia , BERT , OPT , Llama 2 / Llama 2 base , ViT IC-915 In Pythia-70M and Pythia-160M, individual MLP hidden neurons are activated by multiple irrelevant token combinations (pattern superposition) Pythia IC-916 Pythia models show scale-dependent last-layer averaging barriers: 70M exhibits a barrier of ~13 while 410M shows ~1 Pythia IC-917 ViT-S models show early-layer sensitivity to layer-wise averaging, with the averaging direction being far more disruptive than random perturbations of the same norm ViT IC-925 MultiBERTs exhibits the same phase transition pattern (UAS spike, loss drop, BLIMP improvement) as the authors' own BERT-base training MultiBERTs IC-926 Across 25 MultiBERTs seeds, UAS does not correlate with MLM test loss or BLIMP accuracy MultiBERTs IC-927 Human ciphers (ASCII, Unicode, Caesar, Morse) bypass the safety alignment of GPT-4 and GPT-3.5-turbo, with more powerful models producing more unsafe responses GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Falcon , Llama 2 / Llama 2 base , GPT-3 / GPT base IC-928 SelfCipher (a role-play prompt without explicit cipher rules) evokes a 'secret cipher' in LLMs, achieving high unsafety rates that outperform most human ciphers GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 , Falcon , Llama 2 / Llama 2 base , GPT-3 / GPT base IC-929 Simulated character-level ciphers that never appear in pretraining data cannot bypass safety alignment even with 10+ demonstrations GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-3.5 / ChatGPT-3.5 IC-930 StackLLama, when used as a reward model, achieves near-random consistency on contrast instructions for the Stack Exchange task StackLLaMA IC-931 GPT-4 achieves approximately 95% accuracy on contrast instructions, far exceeding human performance without tools GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report IC-932 Pre-trained DNN object detectors show a sharp falloff in peripheral detection performance with increasing eccentricity, degrading to near-chance by 20°, while human performance degrades gradually DINO-FocalNet-Large , Swin Transformer , DETR-R50 , RetinaNet-R50 , FoveaBox , Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 IC-933 DNN object detectors do not exhibit the same sensitivity to image clutter as humans in peripheral object detection DINO-FocalNet-Large , Swin Transformer , DETR-R50 , RetinaNet-R50 , FoveaBox , Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 IC-934 DNN object detectors do not exhibit the same object size effect as humans in peripheral object detection DINO-FocalNet-Large , Swin Transformer , DETR-R50 , RetinaNet-R50 , FoveaBox , Faster R-CNN / Faster R-CNN R50 / Faster R-CNN X101 IC-935 In pre-trained ViT, query vector Kruskal rank reaches the context size only after one self-attention layer, while general position fails at all depths ViT IC-936 GPT-2's learned positional encodings cause context vectors to lose linear independence after one layer, whereas BERT's sinusoidal encodings preserve it GPT-2 , BERT IC-939 CLIP reward landscapes are well-shaped for photorealistic environments but poorly shaped for abstract renderings CLIP / CLIP-ViT (LC) IC-940 CLIP can specify 5 of 8 complex humanoid tasks from single-sentence prompts, failing on tasks requiring discrimination of subtle body-pose differences CLIP / CLIP-ViT (LC) IC-941 CLIP reward model quality scales with model size, with a sharp phase transition between ViT-H/14 and ViT-BigG/14 for the humanoid kneeling task CLIP / CLIP-ViT (LC) IC-945 LLaMA-2, MPT, Falcon, Pythia, and BERT-base-uncased allocate disproportionate attention to initial tokens regardless of their semantic content Llama 2 / Llama 2 base , MPT , Falcon , Pythia , BERT IC-946 LLaMA-2-7B, MPT-7B, Falcon-7B, and Pythia-12B do not consistently improve in perplexity as the StreamingLLM cache size increases Llama 2 / Llama 2 base , MPT , Falcon , Pythia IC-951 The programmatic space achieves behavior-similarity values comparable to the LEAPS latent space, indicating that optimizing the behavior loss alone does not produce a more search-conducive space LEAPS IC-955 All baseline knowledge tracing models either lack significant correlation or negatively predict causal support from their inferred prerequisite graphs HLR , DKTF , AKT , GKT , QIKT IC-959 LLaMA and OPT-1.3B (and Aquila-7B) encode more similar interaction primitives than smaller models such as BERT-base and BERT-large LLaMA , OPT , Aquila-7B IC-981 ImageBind's indirect alignment through images degrades zero-shot performance on non-visual modalities and prevents emergent cross-modal retrieval ImageBind IC-982 In Stable Diffusion's UNET, visual attribute knowledge is distributed across multiple components with attribute-specific patterns, concentrated more in the up-block, and cross-attention layers are not the primary causal states Stable Diffusion IC-983 In Stable Diffusion's CLIP text-encoder, knowledge about all visual attributes is localized to a single causal state: the first self-attention layer at the last subject token Stable Diffusion , CLIP / CLIP-ViT (LC) IC-984 CLIP ViT-B/16's representation space does not reliably preserve semantic similarity as measured by shared image tags CLIP / CLIP-ViT (LC) IC-985 LLaMA-2-7B plateaus in ICL accuracy and fails to override semantic priors when in-context labels are flipped on a simple happy/sad classification task Llama 2 / Llama 2 base IC-986 Most LLMs lack tool usage awareness, with only ChatGPT exceeding 70% F1 in zero-shot evaluation ChatGPT , ChatGLM2 , Llama 2 / Llama 2 base , Vicuna , Koala IC-987 When the correct tool is absent from the candidate list, most LLMs hallucinate a tool rather than returning 'none' ChatGPT , ChatGLM2 , Llama 2 / Llama 2 base , Vicuna , Koala IC-988 LLMs show large gaps in multi-tool selection and over-rely on the number of tools specified in the prompt ChatGPT , ChatGLM2 , Llama 2 / Llama 2 base , Vicuna , Koala IC-989 Tool selection CSR degrades as the candidate tool list grows from 5 to 15 tools, and performance varies by user scenario ChatGPT , ChatGLM2 , Llama 2 / Llama 2 base , Vicuna , Koala IC-990 ResNet-50 and DenseNet-101 exhibit a higher mean-to-variance ratio in penultimate pre-ReLU activations for in-distribution samples than for out-of-distribution samples, and the activation-based scaling factor is well-separated between ID and OOD ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , DenseNet / DenseNet-101 IC-991 LLaMA-2, Falcon-7B, and GPT-3.5-Turbo exhibit large performance spread (up to 76 accuracy points) across semantically equivalent prompt formats, and model comparison rankings are frequently reversed by format choice Llama 2 / Llama 2 base , Falcon , GPT-3.5 / ChatGPT-3.5 IC-992 LLaMA-2-7B's last hidden layer encodes the prompt format with high identifiability, and the separability of format embeddings in the top two principal components correlates with performance spread Llama 2 / Llama 2 base IC-996 GPT-3.5 achieves 73.5% zero-shot accuracy on OGBN-ARXIV 40-class node classification and 73.56% on TAPe-ARXIV23, a dataset of papers published after its knowledge cutoff GPT-3.5 / ChatGPT-3.5 IC-997 LLaMA-2-13B-Chat achieves 44.23% zero-shot accuracy on OGBN-ARXIV, substantially below GPT-3.5's 73.5% Llama 2 / Llama 2 base IC-998 GPT-3.5's zero-shot accuracy on OGBN-ARXIV depends on the position of the title relative to the abstract in the prompt: 0.720 when abstract precedes title, 0.695 when title precedes abstract GPT-3.5 / ChatGPT-3.5 IC-999 FLUX1 and Stable Diffusion 3.5 exhibit high local dependency ratio and produce text hallucinations when generating text content FLUX / FLUX1 , Stable Diffusion SY-001 Objects outside the patient, gown snaps and ECG electrodes, drive Sybil's risk predictions Sybil SY-002 Sybil responds more weakly to nodules near the pleura, where adenocarcinoma tends to appear Sybil SY-003 Sybil processes pulmonary nodules almost additively, with limited pairwise interactions Sybil TM-001 A mean-difference vector between before and after image pairs acts as a transferable concept vector TerraMind TM-002 The mean-difference concept vector scores higher than trained classifiers under cosine similarity TerraMind TM-003 Optical-to-SAR generation errors cluster spatially, but only five regions survive FDR control TerraMind TM-004 Surface composition, not the acquisition time gap, correlates with optical-to-SAR reconstruction error TerraMind TM-005 Flooded vegetation gives the lowest reconstruction error of any land-cover class, not the highest TerraMind TM-006 Coordinates are recoverable from TerraMind's frozen features, latitude more accurately than longitude TerraMind TM-007 Coordinates are linearly decodable only in the larger TerraMind variants TerraMind TM-008 TerraMind embeddings separate hemispheres rather than climate zones TerraMind TM-009 A weak seasonal shift in the embeddings matches climatic intuition but is attributed to geography TerraMind TM-010 Fixed corner patches dominate Integrated Gradients maps regardless of image content TerraMind TM-011 Spatial coherence of latent feature planes increases with encoder depth TerraMind TM-012 Longitude and time-of-year planes become circular in deeper encoder blocks TerraMind TM-013 Intervening on the altitude plane raises generated terrain by almost 500 metres TerraMind TM-014 Per-edge latent distance along shortest paths spikes at physical barriers TerraMind Registries