Light Dark The model relies on a feature that correlates with the target in the data it saw but has no causal connection to it: printed text standing in for the object named in a caption, a chest drain standing in for disease, position on the globe standing in for climate. Shortcuts work until the correlation disappears, which is usually at deployment. The older name for the same thing is Clever Hans behaviour, after Lapuschkin et al. 2019; both names point here rather than at two nodes. How strongly the reliance was established -- by removing the feature and remeasuring, or by beating the model with a baseline built from the spurious feature alone -- is recorded in evidence_type, not in this definition.
Findings FX-003 FIXLIP's strongest interaction in one CLIP example links doll to an image patch reading dollar IC-047 Tulu-2-13B's entity-attribute binding partially relies on token order as a shortcut, degrading in nested orderings where order and semantic binding conflict IC-084 Safety alignment in Llama-2-7b-chat and Gemma-7b-1.1-it is shallow, with the KL divergence from the base model concentrated in the first few output tokens, making the models vulnerable to prefilling attacks IC-085 Unaligned base models Llama-2-7b and Gemma-7b produce predominantly safe continuations when prefilled with refusal prefixes, demonstrating a pre-existing safety shortcut IC-1016 ImageNet-pretrained ResNet-50 backbone exhibits shortcut bias toward background features when adapted to bird classification via a new readout layer IC-1078 LLaMA models (7B through 65B) exhibit gender bias in language generation, coreference resolution, and sentence likelihood, with stereotypical associations driving predictions IC-1139 Stable Diffusion encodes social biases in its internal concept representations that are not always visually apparent IC-1153 Intentionally constructed spurious token connections in MMLU demonstrations misdirect LLaMA-65B's in-context learning toward specific answer choices IC-118 Hallucination heads in LLaVA-7B and MiniGPT-4 allocate 4.75x more attention to text tokens than image tokens, and this pattern is inherited from the base language model IC-1206 GPT-2-XL, GPT-J, Falcon-7B, Llama-2-7B, and Llama-2-13B are vulnerable to backdoor injection via lightweight parameter editing with only 15 samples, achieving near-100% attack success rate while preserving clean performance IC-1244 GPT-4's non-zero scores on GAIA web browsing questions are largely due to memorization of intermediate information from training data rather than actual web browsing IC-131 GPT-4 Turbo, GPT-3.5 Turbo, Llama3-8B, Qwen-7B, and iFlytekSpark-13B over-rely on the strong reminder 'the answer is' in prompts as a shortcut, with accuracy dropping sharply when the cue is a random answer rather than the ground truth IC-132 GPT-4 Turbo, GPT-3.5 Turbo, Llama3-8B, Qwen-7B, and iFlytekSpark-13B trust authority roles (teacher/judge) more than peer roles (classmate/lawyer) when the cue information is the correct answer IC-1390 Proficiency tests reveal that a substantial portion of correct main-test predictions by VidLMs and ILMs are spurious rather than reflecting robust understanding IC-147 CLIP's OOD performance on rendition domains is largely an artifact of domain contamination in its web-scale training data IC-1518 Domain finetuning of LLaMA 2 7B, LLaMA 2 13B, and GPT-2 XL on PubMed causes topic and style priors to shift dramatically, accounting for the majority of the probability change, while factual knowledge learning contributes only a small fraction IC-1519 Topic and style biases in LLaMA 2 7B are learned like simple features (rapidly, with minimal capacity, concentrated at the first few tokens, magnified by learning rate) while factual knowledge is learned like complex features (slowly, requiring significant capacity, uniformly across positions, unaffected by learning rate) IC-152 CLIP's polysemantic neurons encode spurious correlations between unrelated concepts that can be exploited to generate adversarial misclassifications IC-153 ResNet50 relies on flower petals and green background features as shortcuts when classifying bee images IC-155 All 13 evaluated MLLMs perform at or near random guessing on MediConfusion, with confusion scores often exceeding 90%, indicating they cannot distinguish visually dissimilar radiology image pairs IC-161 Linear probes on Pythia-70m and Gemma-2-2b trained on the ambiguous Bias in Bios set rely on gender as a spurious feature, with gender accuracy far exceeding profession accuracy IC-169 BT-based, DPO-based reward models, and GPT-4 as judge all exhibit significant length bias, with their scores correlating with output length rather than quality IC-193 The OpenCLIP ResNet-50 model trained on CC12M contains an unintentional backdoor from birthday cake images in CC3M, achieving 98.92% attack success rate IC-198 Safety-aligned LLMs (GPT-4, GPT-3.5, Gemma2-27b, GPT-4o, Gemma2-9b, Qwen2.5-72b, Mistral-7b, Mixtral-8x22b) are vulnerable to natural prompts semantically related to toxic seed prompts, with attack success rates of 82-99% IC-209 LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferences IC-259 Fine-tuned DNNs approach human accuracy on VPT-basic but fail on VPT-strategy, revealing reliance on a brittle feature-based shortcut (object size and location) rather than line-of-sight estimation IC-278 Claude-3.5 Sonnet and GPT-4o exhibit a memorization failure mode, outputting the same answer regardless of visual parameter changes in the problem IC-396 Closed-source LLMs (GPT, Claude) rely more on deep structure than open-source LLMs (Llama, Mistral), and open-source models' surface sensitivity decreases with model scale IC-411 Llama2-7B, Llama3-8B, and Mistral-v0.3-7B do not reason with edited knowledge in multi-hop questions, as editing methods mostly underperform pre-edit portability scores IC-416 Natural language prompts can steer the texture/shape bias in VLMs in both directions without significantly affecting accuracy, with texture-biased prompts more effective than shape-biased ones; this steering also generalizes to low/high-frequency bias. IC-456 VLM decoders achieve near-random accuracy on VALSE image-sentence alignment while pairwise accuracy is much higher, indicating reliance on linguistic priors IC-457 All four tested VLM decoders are heavily text-centric when generating answers, with text modality contributing 85-97% of the prediction signal IC-485 LLMs show a significant performance gap between Wikipedia-based factual multi-hop QA and counterfactual multi-hop QA, indicating reliance on memorized knowledge rather than reasoning from context IC-498 Multimodal foundation models exhibit severe group unfairness, with race and age biases more pronounced than gender bias in text-to-image models while gender bias is stronger in image-to-text models IC-587 Real-world knowledge acts as a shortcut in LLM relational reasoning, causing worse-than-chance performance on logically valid but factually incongruent statements IC-679 CLIP relies on background/location as a spurious cue for bird classification, and ablating geolocation heads improves worst-group accuracy by 25.2% IC-685 CLIP ViT-B/32 with a linear probe relies on gender as a spurious correlation for hair color, achieving only 15.85% accuracy on female gray hair IC-742 ResNet-50-BN on Waterbirds relies on background as a spurious feature for classification, and this shortcut is invisible to entropy-based confidence metrics IC-835 CLIP zero-shot predictions exhibit large worst-group accuracy gaps due to spurious correlations on Waterbirds and CelebA IC-836 CLIP zero-shot predictions exhibit demographic bias on FairFace when using attribute-unrelated text prompts IC-859 LLMs rely on learned priors about graph properties (cycles exist, edges are absent) rather than analyzing the specific graph structure, causing below-majority-baseline performance and extreme structure-dependent accuracy SY-001 Objects outside the patient, gown snaps and ECG electrodes, drive Sybil's risk predictions TM-008 TerraMind embeddings separate hemispheres rather than climate zones TM-009 A weak seasonal shift in the embeddings matches climatic intuition but is attributed to geography