Light Dark A reproducible condition under which the model produces a wrong or degraded output. The finding names the condition, not a single bad example, so it has to say what the condition is and over how many cases it was checked.
Findings IC-001 Vision-language models perform near chance on the NL-Eye visual abductive reasoning benchmark IC-002 Even when VLMs select the correct hypothesis, their explanations are often invalid or unhelpful IC-005 GPT-4o underperforms AHA and other VLMs in detecting and reasoning about robotic manipulation failures across multiple datasets. IC-008 Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances IC-021 Vision-language adaptation degrades safety in Llama-2-chat-7b even when training data is filtered for safety IC-025 LMMs exhibit poor fine-grained perception in locating individual characters on original oracle bones IC-026 LLMs can assist in OB rejoining by identifying rejoinable fragments with moderate accuracy, but are not yet truly usable IC-027 LMM performance in deciphering oracle bone inscriptions is comparable to untrained humans for common characters but declines for rarer and structurally complex characters IC-029 Large language models show conformity to group answers in multi-agent interactions IC-032 Off-policy DPO causes a squeezing effect in LLMs where probability mass shifts to the most confident token, explaining degenerate repetition IC-040 Information imbalance triggers both the modality gap and object bias in contrastive VLMs IC-046 Tulu-2-13B exhibits gender bias in both its internal binding representation and its outputs, with the output-level bias being stronger than the representation-level bias IC-048 Existing video LLMs (TimeChat, VTG-LLM, Momentor, Hawkeye) show limited zero-shot video temporal grounding capability and struggle to improve with fine-tuning IC-049 GPT-4o shows strong performance on some E.T.Bench event-level tasks (RVQ: 57.7, VHD: 56.9) but very weak performance on others (EPM: 4.5, TAL: 20.0) IC-050 GPT-4o achieves 53.33 overall on Event-Bench with strong event description (57.50) and counter reasoning (63.44) but weaker episodic reasoning (37.33) IC-051 Qwen2-VL (7B) achieves 0.0 on DVC, DVC SLC, and TEM tasks on E.T.Bench IC-069 EVA-CLIP's dense patch features are semantically contaminated by surrounding context, degrading their spatial quality IC-070 Region-language alignment fine-tuning degrades EVA-CLIP's spatial awareness as measured by unsupervised segmentation IC-071 DINOv2's dense features are dominated by global context, impairing fine-grained spatial detail IC-072 LLM agents of varying scales exhibit a failure mode on web automation tasks when processing raw, complex web page observations, with the penalty being more severe for smaller models IC-073 Released LLMs (GPT-4o, Llama-3.1-70B, Qwen2-7B, etc.) show limited workflow orchestration capability that degrades as workflow complexity increases IC-075 GPT-4o-mini and Qwen2.5-72B achieve low precision and recall when used as API retrievers for workflow orchestration IC-078 GPT-4's self-verification loop causes performance collapse due to high false negative rates in binary verification IC-079 GPT-4's free-form critique generation is unreliable, containing hallucinated edges, vertex colors, and precondition states IC-080 GPT-4's performance is largely insensitive to the content of feedback; simple re-prompting with a sound verifier (sampling) matches or exceeds detailed critique IC-084 Safety alignment in Llama-2-7b-chat and Gemma-7b-1.1-it is shallow, with the KL divergence from the base model concentrated in the first few output tokens, making the models vulnerable to prefilling attacks IC-086 Fine-tuning Llama-2-7b-chat on 100 harmful examples for 6 gradient steps increases the attack success rate from 1.5% to 87.9%, with per-token dynamics showing the distributional change concentrated in the first few tokens IC-092 All evaluated LLMs show consistent F1 degradation to at most 0.60 when two or more events match a retrieval cue IC-093 No evaluated LLM achieves perfect confabulation avoidance on questions about non-existent events IC-094 Episodic recall accuracy degrades systematically from content cues to space cues to time cues across all evaluated LLMs IC-095 Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological ordering IC-098 LegalBERT performs below the constant classifier baseline on CASELAWQA due to its 512-token context window IC-099 GPT-4 and Claude 3 Opus can be prompted to selectively underperform on WMDP while maintaining general performance on MMLU and CSQA IC-1007 LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback IC-101 GPT-4 and Claude 3 struggle to emulate a lower capability profile (high school freshman level) via zero-shot prompting, with only moderate improvement from chain-of-thought prompting IC-1011 OpenAI CLIP loses approximately 8% zero-shot retrieval accuracy on 2021–2022 data compared to OpenCLIP models trained on data through 2022, while standard benchmarks show no such gap IC-1012 Connected regions in the latent space of Stable Diffusion v2.1, v1.5, and GLIDE produce distorted images independent of the text prompt IC-1013 Latent samples in Stable Diffusion v2.1, v1.5, and GLIDE can produce images of associated backgrounds rather than the key object, with failure rates of 2.7%, 9.2%, and 50.5% under random sampling respectively IC-1014 A single adversarial token embedding appended to any input prompt overwrites the prompt in Stable Diffusion v2.1 to generate a target object, with CLIP similarity to the original prompt (0.742) remaining higher than to the target (0.546) IC-1015 GPT-J and 10 other LLMs exhibit overthinking: calibrated accuracy given incorrect few-shot demonstrations peaks at a critical layer then declines, and ablating 5 false induction heads in late layers reduces the accuracy gap by 38.9% on average IC-102 GPT-4o's spatial understanding degrades when depth maps are provided as additional input on SpatialBench IC-1035 Reprover achieves 0% accuracy on sorry theorems in advanced mathematics repositories (PFR, Hairy Ball Theorem, Coxeter) while proving basic theorems in other repositories IC-1044 Counterfactual GNN explainers produce statistically infeasible recourses that violate topological constraints in molecular datasets IC-1050 Released models generate patches that are less than half the length of gold solutions and rarely edit more than one file IC-1053 Instruction tuning suppresses in-context learning in LLaMA, Vicuna, and OPT-IML, with the suppression being largest for English prompts and partially recoverable via translation to other languages IC-1054 Code fine-tuning degrades English natural language reasoning in Code LLaMA relative to LLaMA-2, but the effect is negligible or slightly positive in French, Spanish, and German IC-1055 Safety fine-tuning suppresses harmful content generation in ChatGPT relative to GPT-3.5, but the suppression is substantially weaker for non-English prompts IC-1086 RS LDS fails to increase its number of active states when the underlying dynamics change non-stationarily IC-1087 SLDS produces poor dynamical accuracy on the NASCAR task because it lacks recurrent switching IC-1089 ICL in LLaMA, LLaMA-2, and Falcon models cannot fully overcome pre-training label preferences when in-context labels are flipped IC-109 GPT-3.5-turbo and GPT-4 produce cycles in inferred causal graphs when using pairwise prompts, with cycle counts growing sharply on larger graphs IC-112 Released LLMs show a reproducible failure mode where numerical task accuracy degrades sharply as input digit length increases IC-1123 All seven published concept erasure methods applied to Stable Diffusion 1.4 can be circumvented by learned word embeddings, demonstrating that targeted concepts are input-filtered rather than truly removed from the model IC-113 Released LLMs show a reproducible failure mode where accuracy on fraction and scientific notation tasks falls below 20% even for the shortest inputs IC-114 Released LLMs cannot reliably identify a specific digit in a number as the number's length increases, with GPT-4o achieving only 20% on get-digit in the xl range IC-1144 GPT-4 and GPT-3.5 produce high rates of irrelevant (fabricated) books when answering constraint queries from parametric knowledge, with a sharp phase transition at low author popularity IC-1145 Providing complete context eliminates irrelevance but does not fix constraint satisfaction for GPT-4 or GPT-3.5 IC-1146 Self-context (self-retrieval) chain-of-thought increases the rate of fabricated books compared to no-context for both GPT-4 and GPT-3.5 IC-1149 ChatGPT and Llama-2-7b-chat fail to recognize unanswerable questions on SQuAD 2.0, with Llama-2-7b-chat scoring only 3.72% accuracy on no-answer questions IC-1150 ChatGPT and Llama-2-7b-chat underperform humans by 20 and 31 points respectively on out-of-distribution NLU tasks in GLUE-X IC-1153 Intentionally constructed spurious token connections in MMLU demonstrations misdirect LLaMA-65B's in-context learning toward specific answer choices IC-1157 GPT-4 and other LMs show a large gap between rule induction and rule application, with task accuracy dropping to near zero on MiniScan when the LM itself applies its own proposed rules IC-1158 GPT-4 and other LMs are brittle to noisy exemplars and unfamiliar output representations, with performance degrading sharply even under minimal perturbation IC-1170 GPT-3.5, Llama2, PaLM2, and GPT-4 are susceptible to a CoT-prompting backdoor attack (BadChain) on complex reasoning tasks, with stronger reasoning models showing higher attack success rates IC-1175 ChatGPT can be prompted to generate misinformation with near-perfect success for implicit methods but is largely resistant to explicit misinformation requests IC-1177 LLM-generated misinformation is harder for LLM detectors to detect than human-written misinformation with the same semantics IC-1193 Vicuna and Alpaca achieve 0% pass rate on all ToolBench tool-use instructions, while GPT-4 and ChatGPT reach 71.1% and 64.8% with DFSDT, revealing a wide capability gap in tool use among released LLMs IC-1194 GPT-3.5-turbo, GPT-4, and GPT-3.5-turbo-0613 exhibit 50-58% inconsistency between their ratings and rankings feedback on the same response pairs IC-1195 Model substitution adversarial attack reduces TreeRing AUROC to 0.14 at ε=2/255 and StegaStamp AUROC to 0.492 at ε=12/255 IC-1196 Blending a watermarked noise image with a clean image causes watermark detectors to falsely flag clean images as watermarked IC-1199 GPT-2 next-token distributions contain correctable tail errors from the softmax bottleneck that degrade generation quality under low-entropy sampling, with basis-aware threshold sampling improving MAUVE across all four sizes IC-1202 Stable Diffusion 1.5 and 2.1 exhibit a cropping failure mode where synthesized objects are cut off at image boundaries IC-1205 GPT-3 models (ada, curie, davinci) achieve near-zero accuracy on zero-shot arithmetic tasks but learn them rapidly with 1000 fine-tuning samples IC-1206 GPT-2-XL, GPT-J, Falcon-7B, Llama-2-7B, and Llama-2-13B are vulnerable to backdoor injection via lightweight parameter editing with only 15 samples, achieving near-100% attack success rate while preserving clean performance IC-1217 LLaMA 65B's token-probability readout fails to capture human decision-making, producing near-chance NLL and no human-like exploration behavior IC-1232 GPT-2 XL and GPT-J exhibit knowledge conflict when subjected to reverse and composite knowledge edits, with ROME and MEMIT showing near-total failure on reverse edits IC-1233 GPT-2 XL and GPT-J exhibit irreversible knowledge distortion after round-editing, with the effect being more severe when the edit target is semantically distant from the true labels IC-1237 Llama-2-7b's pre-existing translation knowledge is diluted by large amounts of parallel data, causing COMET to decline after 100k examples IC-1238 Llama-2-13b produces off-target non-translation outputs in zero-shot English-to-foreign-language translation IC-124 Direct comparison of CLIP image embeddings with CLAP audio embeddings achieves near-chance retrieval, while logsumexp bridging through the shared language modality recovers 62% recall@10 on AudioSet IC-1243 GPT-4 and GPT-4 Turbo achieve near-zero scores on GAIA level 3 and single-digit to low-double-digit scores on levels 1-2, compared to 87-94% for human annotators IC-1249 TD-MPC exhibits training instability and performance degradation when the planning horizon is set to 20 time steps IC-1256 MPT-7B-Chat produces non-committal responses rather than proper refusals on unsafe instructions IC-1257 Guanaco acknowledges the illegality of requested actions but still provides the harmful information IC-1264 LLMs are overconfident when verbalizing confidence, with values concentrated in 80–100% and multiples of 5, yielding high ECE across all five tested models IC-1267 LLMs' alignment with human privacy judgments drops sharply as contextual complexity increases from tier 1 to tier 3 IC-1268 LLMs leak private information in theory-of-mind scenarios even when explicitly instructed to preserve privacy IC-1269 LLMs leak secrets to inappropriate recipients in meeting summarization and action-item generation tasks IC-1270 Chain-of-thought prompting does not mitigate privacy leakage in GPT-4 or ChatGPT IC-128 Released VLMs exhibit sycophancy, agreeing with incorrect user opinions while ignoring visual evidence, with LLaVA-1.5 showing the highest rate (94.6%) and InternLM-XComposer2-VL-1.8B the lowest (28.8%) IC-1284 Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on as few as 10 explicitly harmful examples removes their safety alignment, raising harmfulness rates to 80-92% IC-1285 Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on 10 implicitly harmful identity-shifting examples (containing no toxic content) jailbreaks their safety alignment IC-1286 Fine-tuning GPT-3.5 Turbo and Llama-2-7B-Chat on benign utility-oriented datasets (Alpaca, Dolly, LLaVA-Instruct) degrades their safety alignment without any malicious intent IC-1287 A backdoor can be implanted in GPT-3.5 Turbo via fine-tuning that is undetectable by standard safety auditing: the model appears safe on plain prompts but fulfills harmful instructions when a 3-word trigger is appended IC-1288 GPT-4, ChatGPT, and GPT-4V fail to close the human-machine gap on Bongard-OpenWorld, with InstructBLIP captions differentially degrading ChatGPT while improving GPT-4 IC-1289 OpenFlamingo and Otter achieve near-chance accuracy on Bongard-OpenWorld, indicating inability to perform multi-image reasoning IC-1290 CLIP, DINO, and DINOv2 as zero-shot natural baselines score below the 50% chance level on Bongard-OpenWorld due to adversarial query selection IC-1307 Released LLMs (CodeLlama 7B/13B/34B, GPT-3.5, GPT-4) achieve limited code-optimization speedups with standard prompting, with the best baseline (GPT-3.5 CoT) reaching only 1.60x versus the 3.66x human reference IC-1309 GPT-4-0613 exhibits reduced output diversity relative to GPT-3.5: it outperforms on best@1 but underperforms on best@8 under CoT prompting IC-1327 GPT-4, GPT-3.5, Llama2, and Vicuna models underperform human annotators on multistep soft reasoning in natural language narratives, with smaller models scoring near random chance IC-1332 Vicuna-v1.5 and CodeLlama-34b-instruct produce format-breaking artifacts (escaped underscores, [python] tags) in 30-100% of code instances due to training data contamination IC-1336 MobileNetV2 (PyTorch pre-trained on ImageNet) exhibits a failure mode under global unstructured L1 pruning at the Pareto-optimal point, with its high kurtosis of kurtoses (64.40) causing very-low-magnitude layers to be entirely pruned and disconnect the network IC-134 Mistral Large V2 as a judge on SummEval coherence systematically avoids extreme ratings (1 and 5), while human annotators assign over 24% of items a median rating of 5 IC-1342 Assigning socio-demographic personas to LLMs causes significant reasoning performance degradation across all four models studied, manifesting as both explicit abstentions and implicit reasoning errors IC-1355 Instruction-tuned VLMs fail to follow multiple-choice format in reasoning questions, with InstructBLIP frequently returning blank responses IC-1374 Counting difficulty in constrained generation increases with text level and constraint strictness, with exact sentence-level character counts being the hardest condition for all five models IC-138 Trojan backdoored Llama-2-7B models and Vicuna-7B-v1.5 exhibit the probe concatenate effect, where concatenating a triggered or jailbroken sample with a harmful probe significantly shifts the model's output distribution away from refusal IC-1389 Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMA IC-139 3D Gaussian Splatting and its variants (Scaffold-GS, Mip-Splatting) are vulnerable to computation cost attacks via data poisoning, with peak GPU memory increasing up to 21.93x and training time up to 4.97x under unconstrained perturbation IC-1391 SLD concept removal variants and SD with negative prompts are bypassable by Ring-a-Bell adversarial prompts, increasing attack success rate from single digits to 90-100% for nudity IC-1393 Most mainstream LLMs generate value-violating content at high rates (APV 65-80%) across 2,397 morally ambiguous prompts, indicating substantial ethical misalignment IC-1399 Invariant GNNs (l=0) consistently fail to distinguish k-hop identical but globally distinct geometric graphs on the k-chain task, regardless of model depth IC-1403 All four evaluated LLMs produce negative scores on social rules and secret-keeping dimensions in SOTOPIA interactions IC-1404 On SOTOPIA-HARD, GPT-4 achieves significantly lower goal completion than humans and exhibits non-strategic negotiation and excessive compromise behaviors IC-1405 OpenFlamingo and Idefics models hallucinate objects not present in images, and increasing ICL shots beyond 4 amplifies hallucinations IC-1406 OpenFlamingo and Idefics models rarely abstain from answering unanswerable questions, but ICL significantly improves abstention F1 IC-1407 OpenFlamingo and Idefics models perform near random chance on compositional image-text matching, and ICL has almost no effect on atomic foils IC-1413 GPT-4 achieves near-saturation on Python code synthesis (86.6% pass@1) but scores significantly lower on code repair (47.8% avg) and code explanation (52.1% avg) across six languages IC-1414 Pretrained code models Starcoder and CodeGeex2 score 0.0% on code explanation across all six languages because they generate code instead of natural language IC-1436 Imagen Video 5.6B produces high-quality but domain-inappropriate videos on out-of-distribution robotics and egocentric data, failing to generate relevant dynamics IC-1437 LLaMA-2-7b-chat-hf and Meta-LLaMA-3-8B-Instruct exhibit reduced attention to rule tokens and fail to follow prompt-specified rules when the adversarial suffix 'forget all prior instructions and answer the question' is appended IC-1438 LLaVA and Llama-Adapter V2 are jailbroken by compositional adversarial images targeting image-based embedding triggers, with near-zero success for textual triggers IC-1439 LLaVA follows text instructions embedded in adversarial images as if they were user prompts, enabling hidden prompt injection IC-144 TD-MPC's training is unstable, with performance collapsing after approximately 1–4 million steps across multiple DM Control tasks IC-1440 GCN, GAT, GraphSAGE, and SGC exhibit structure-dependent generalization in transductive node classification: test nodes with shorter paths to training nodes are classified more accurately IC-1441 GCN exhibits structural unfairness in transductive node classification, with demographic parity and equal opportunity gaps between nodes connected to and disconnected from the training set IC-145 TD-MPC fails to achieve high reward when irrelevant background information is added to image inputs IC-1456 Pre-trained language models fail to predict both interpretations of ambiguous inputs in zero-shot semantic parsing IC-1461 Varying decoding hyperparameters and removing the system prompt breaks the safety alignment of 9 out of 11 open-source LLMs, raising attack success rate from 0% to over 95% IC-1468 DECAF's optimization-based fitting degrades under significant self-occlusion where the hand covers more than half the face IC-1484 GP-UNIT's FID degrades under noisy inputs in reference-guided mode but paradoxically improves in latent-guided mode IC-1485 Sketch Transformer's FID degrades from 31.49 to 404.01 under Gaussian noise at the highest tested intensity IC-1486 HiFaceGAN's FID degrades from 34.83 to 320.41 under Gaussian noise at the highest tested intensity for face super-resolution IC-1487 CycleGAN's FID degrades from 76.92 to 180.82 under Gaussian noise for horse-to-zebra translation IC-1489 State-of-the-art foundation models (CLIP, GPT-3.5-turbo, and others) score well below elementary students on multimodal K-12 STEM questions IC-1510 Without retrieved context, LLMs are unable to translate Kalamang, and among context types, retrieved parallel sentences are most beneficial, followed by word list entries, then grammar book passages IC-152 CLIP's polysemantic neurons encode spurious correlations between unrelated concepts that can be exploited to generate adversarial misclassifications IC-1523 Five AI assistants (Claude-1.3, Claude-2.0, GPT-3.5-turbo, GPT-4, Llama-2-70B-Chat) consistently exhibit sycophancy across four varied free-form text-generation tasks IC-154 CLIP ViT-L/14 text embeddings fail to capture fine-grained visual class similarities, ranking rottweiler and doberman at position 828 behind unrelated pairs IC-1543 VGG19, ResNet50, ViT-Base, and DeiT-Base (ImageNet pretrained) achieve near-zero accuracy under query-based black-box attacks with 1000–10000 queries IC-1549 All 28 evaluated LMs exhibit gender bias on non-stereotypical sentence pairs, with fairness scores between 9% and 41% IC-155 All 13 evaluated MLLMs perform at or near random guessing on MediConfusion, with confusion scores often exceeding 90%, indicating they cannot distinguish visually dissimilar radiology image pairs IC-1550 All evaluated LMs systematically prefer male pronoun completions in the non-stereotypical portions of Winobias and Winogender, with margins exceeding 40% IC-1555 GPT-J's internal representations contain correct factual knowledge even when the model outputs falsehoods under repetition or instruction distraction prompts IC-1558 Code LLaMA 13B maintains 99.4% passkey retrieval at 128k context despite perplexity rising from 2.37 to 2.54 between 98304 and 131072 tokens IC-156 Gemini models show substantially lower confusion scores than other MLLMs yet still perform at or near random guessing, suggesting their bottleneck is medical knowledge or reasoning rather than visual encoding IC-157 GPT-4o's MediConfusion performance is robust to prompt format while InstructBLIP is highly sensitive and LLaVA-Med fails completely on multiple-choice evaluation IC-1578 XLM-R-XL without instruction tuning produces [pad] tokens and fails to complete instruction-following tasks IC-1579 ADM's noise prediction network exhibits exposure bias: during iterative sampling the l2-norm of its ε prediction is systematically larger than during training, and the sampling distribution variance exceeds the training variance with error accumulating toward the end of the chain IC-158 Fine-tuning LLaVA-Med on MediConfusion training pairs cannot achieve 100% training accuracy, indicating the vision encoder's embeddings are fundamentally ambiguous for the confusing pairs IC-1580 The official Llama 2-7B checkpoint fails to generate valid numerical responses for 3D-dependent molecular properties, with a valid answer rate of only 23% for SCF energy IC-1584 LLaMA-7B and GPT-J-6B fail to interpret textual emphasis markers, with marked prompting degrading performance substantially IC-1599 Self-repair at equivalent compute budget provides only modest and inconsistent gains over i.i.d. sampling for CodeLlama-13B-Instruct, GPT-3.5, and GPT-4 on HumanEval and APPS IC-1600 Replacing a model's self-generated feedback with a stronger model's feedback consistently improves self-repair beyond both the i.i.d. baseline and the self-repair baseline IC-1601 GPT-4's self-generated feedback is significantly less effective than human programmer feedback for code repair, with the gap widening on harder problems IC-1610 Llama-2-7b-chat underperforms on small molecule editing tasks due to limited domain-specific pretraining IC-1611 Galactica-6.7b fails on protein secondary structure editing tasks, producing hit ratios below random mutation IC-1618 SAM alone has limited generalization for semantic segmentation, producing ambiguous multi-mask outputs without semantic categories IC-163 LLaVA-1.5, LLaVA-Next, and GPT-4V show near-zero accuracy on GUI grounding benchmarks while achieving 50-85 on general image grounding (RefCOCO+), indicating a failure mode specific to GUI grounding scenarios IC-164 Llama-3-8B and Llama-2-7B fail to learn out-of-distribution functions through in-context learning, defaulting to in-distribution predictions IC-170 GPT-4, GPT-3.5, and Claude-3.5-Sonnet rely heavily on parametric knowledge in RAG settings, producing ungrounded responses with high answered ratios and low trust-scores IC-171 ICL prompting produces binary response patterns in released LLMs, with answered ratios collapsing to near 0% or 100% rather than calibrated refusal, making prompting ineffective for RAG groundedness IC-174 RAG reduces model abstention and LLMs hallucinate rather than abstain when the retrieved context is insufficient to answer the query IC-176 LLaMA 3.1 8B Instruct's KGQA accuracy degrades with increasing numbers of retrieved triples, while GPT-4o-mini's accuracy improves, revealing different context-handling capacities IC-177 GPT-4o mini, GPT-4o, and Llama-3-8B all over-rely on incorrect external context, producing wrong answers at high rates when the context conflicts with their internal knowledge IC-179 GPT-4o mini, GPT-4o, and Llama-3-8B all calibrate confidence in their internal answers significantly better than confidence in external contexts IC-184 Mistral-7B, Mistral-7B-Instruct, Llama3-8B, and Llama3-8B-Instruct can internally encode the correct answer while externally generating an incorrect one, with the discrepancy most pronounced for error types where the model shows no external preference for the correct answer IC-188 LLMs show constraint-type-specific performance on system message following, with weaker models exhibiting large variance across constraint categories IC-189 Most LLMs show degraded instruction satisfaction when user instructions conflict with system messages, indicating difficulty in prioritizing system message constraints IC-190 LLMs show progressive degradation in system message constraint following across multi-turn conversations, with dependent conversations degrading faster than parallel ones IC-192 CLIP's contrastive image-text training objective hinders its ability to rank or order images, yielding near-chance performance on ranking tasks in both zero-shot and fine-tuned settings IC-198 Safety-aligned LLMs (GPT-4, GPT-3.5, Gemma2-27b, GPT-4o, Gemma2-9b, Qwen2.5-72b, Mistral-7b, Mixtral-8x22b) are vulnerable to natural prompts semantically related to toxic seed prompts, with attack success rates of 82-99% IC-199 GPT-4o generates natural jailbreak questions from toxic answers without denial, demonstrating an asymmetry in safety training where forward safety (question-to-answer) does not guarantee reverse safety (answer-to-question) IC-202 All six evaluated LLMs achieve very low accuracy on OpenRCA, with no model solving any three-element root cause query IC-203 Gemini 1.5 Pro's RCA-Agent accuracy drops 68.4% when code execution fails, far exceeding the drops for Claude 3.5 (17.9%) and GPT-4o (15.6%) IC-204 GPT-4o performs worse with explicit chain-of-thought prompting than with the original prompt on OpenRCA tasks IC-209 LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferences IC-215 DINOv2's zero-shot attention maps focus on irrelevant foreground objects (vehicles, advertisements) rather than scene structure, degrading its VPR recall on challenging datasets IC-217 AnyLoc's VLAD aggregation, learned unsupervised on the gallery, fails to generalise to out-of-distribution queries with large time gaps or seasonal changes IC-219 The bag-of-heuristics mechanism in LLaMA3-8B fails on certain arithmetic prompts due to insufficient total logit contribution from heuristic neurons, not due to a lack of associated heuristics IC-221 GPT-4o ReAct success rate drops from 47% on synchronous to 11% on asynchronous planning tasks, and all other tested LLMs show equal or worse performance IC-222 GPT-4o ReAct failures are dominated by rule violations (transition function) and goal misinterpretation, with the balance shifting from goal-dominant in synchronous to transition-dominant in asynchronous settings IC-223 GPT-4o ReAct shows poor recovery from failures in asynchronous settings, with 58.6% of failed runs making little to no progress toward the goal and significantly higher repeated transitions than in synchronous settings IC-224 GPT-4o ReAct cannot incorporate stochastic state changes, with success rate on cutting tasks dropping from 56% to 1% when a 33% chance of a cut item reverting to uncut is introduced IC-226 ESM3 (pre-trained, without fine-tuning) shows significantly degraded conformation generation validity at low sampling temperatures (t < 0.5) IC-228 ESM3 (1.4B) fails to capture MD ensemble statistics on the ATLAS benchmark, achieving pairwise RMSD correlation of only 0.08 IC-229 GIN, GCN, and GAT achieve only random-chance accuracy on WL-separable ε-tree graph pairs, exposing a gap between theoretical expressivity and practical separation IC-231 GPT-4-turbo and Claude-3-Haiku show inconsistent adherence to their providers' stated design principles when facing value conflicts in daily-life dilemmas IC-232 System prompts cannot effectively steer GPT-4-turbo's value preferences in moral dilemmas IC-234 SpeechGPT exhibits poor speech-text alignment (ASR-WER 45.00) and degraded response quality in speech-to-speech interaction IC-235 Salmonn and Qwen2-Audio produce responses containing formatted content and redundant explanations that are unsuitable for speech interaction IC-236 Zero-shot Grounding-DINO and Florence-2 show a significant performance gap on referring expression comprehension compared to their fine-tuned versions IC-237 A zero-shot Llama 3 8B, when prompted to select the best bounding box from VLM candidates without fine-tuning, produces results nearly identical to the VLM alone IC-238 LLMs fail to follow user preferences in zero-shot settings, with accuracy below 10% at 10 turns and near zero at 300 turns IC-239 Implicit preference forms (choice-based and persona-driven) are significantly harder for LLMs to follow than explicit preferences at the same context length IC-242 Most LLMs exhibit higher bias ratios in multi-turn dialogues than in single-turn, with bias accumulating across successive turns IC-244 No LLM demonstrates consistently strong fairness across both comprehension-focused and bias-resistance multi-turn tasks; models show complementary failure patterns IC-245 Pretrained LLMs produce duration-dependent outputs that are incompatible with a discrete token interpretation IC-248 Instruction fine-tuning causes context reliance under knowledge conflicts to initially increase then decrease (context-parametric inversion) in Llama2-7B, Pythia-6.9B, and Mistral-7B IC-251 CLIP's overall MMEB performance drops by 29.4% when task-specific instructions are prepended to queries, with classification degrading by 59.3% IC-252 GPT models produce harmful gender stereotypes at higher rates when user names imply a demographic group, with GPT-3.5 Turbo showing the highest rates and open-ended generation tasks most affected IC-254 GPT-4o Mini responses to female-sounding names systematically use simpler, more light-hearted, and less technical language compared to male-sounding names across multiple task domains IC-257 327 DNNs approach or exceed human accuracy on object depth order but are near chance on VPT-basic, while humans show the opposite pattern IC-264 All 18 evaluated LLMs fail to abstain when the provided context lacks the answer, with performance gaps of 13.6% to 68.4% relative to the original context IC-265 Model families show extreme variation in detecting conflicting answers in inconsistent contexts, with phi-3 series at 5.8% average accuracy versus GPT-4 series at 89.35% IC-266 GPT-4o drops from 96.3% closed-book accuracy to 47.5% when given counterfactual context that contradicts its parametric knowledge, far below the 95% human accuracy on the same items IC-267 Adding a 'conflict' instruction to the prompt degrades GPT-4o and Claude 3.5 Sonnet accuracy on normal (answerable, consistent) contexts by 5% and 2% respectively IC-268 ESM-2 and ProGen-2 zero-shot fitness prediction follows an inverted U-shape as a function of wild type sequence likelihood, with both under- and over-preferred sequences degrading performance IC-270 Unsupervised finetuning (evo-tuning) on homologous sequences improves ESM-2 650M zero-shot fitness prediction for low-likelihood wild types but harms high-likelihood ones, with optimal threshold at log-likelihood ε = −1.4 IC-275 Mistral 7B Instruct exhibits a reasoning-type-dependent failure mode where certain problems are exclusively solvable by one non-deductive reasoning type IC-276 All 14 evaluated VLMs show a large gap between average-case and worst-case accuracy on DynaMath variants, with worst-case at or below 50% of average-case, and the failures are systematic rather than random IC-278 Claude-3.5 Sonnet and GPT-4o exhibit a memorization failure mode, outputting the same answer regardless of visual parameter changes in the problem IC-306 Instruction tuning increases the ROUGE score of prompt-injected data extraction by 65.76 on average compared to base models IC-308 Linguistic mutations to unsafe prompts significantly and inconsistently alter safety refusal across models, with persuasion techniques increasing fulfillment by 5-66% and encoding/encryption decreasing it by 15-68% IC-310 Prefilling model responses with 'sure, here is' increases safety fulfillment by 19-58%, and missing prompt template tokens increases fulfillment by 8-30% for Llama-2 and Gemma but not Llama-3 IC-313 All four LLM backbones fail to detect a planted linear trend in incident resolution time when the slope is below 0.1, and detection rates diverge sharply above that threshold IC-316 Multi-turn code generation without CoT degrades performance for smaller Llama models and GPT-4o compared to single-turn repeated sampling under equal compute budgets IC-317 More detailed execution feedback (LDB) induces exploitative behavior in Llama 3.1 models, reducing code diversity and hurting performance at large sample budgets IC-322 Past-tense reformulations of harmful requests bypass refusal training in eight released LLMs, while future-tense reformulations are substantially less effective IC-323 O1-mini and O1-preview reasoning models are vulnerable to past-tense reformulations (84% and 78% ASR) but produce less specific jailbroken outputs than non-reasoning models IC-324 Fine-tuning Gemini Nano 1 on 8 memorization examples causes it to override in-context predictions with in-weight predictions in 2 of 8 cases, while the base model always follows in-context predictions IC-325 A single FFN-layer weight edit (JailbreakEdit) raises jailbreak success rate to 62–87% on Llama-2-7b-chat, Llama-2-13b-chat, Vicuna-7b, and ChatGLM-6b while preserving safety performance and generation quality on non-triggered queries IC-328 Llama-3-8B-Instruct and Mistral-7B-Instruct-v0.3 are susceptible to confidence-elicitation-guided word substitution attacks, with CEAttack outperforming existing hard-label black-box methods IC-329 GPT-4o is more robust to confidence-elicitation-guided word substitution attacks than open-source LLMs, with lower attack success rates and better confidence calibration IC-330 Low local intrinsic dimension (LIDθ) of the learned manifold predicts memorization in Stable Diffusion v1.5, IDDPM, and StyleGAN2-ADA IC-331 In Stable Diffusion v1.5, specific tokens in text prompts drive memorization, and GPT-4-based perturbation of high-attribution tokens reduces SSIM similarity to training images while maintaining CLIP score IC-337 GPT-3.5-based detection accuracy drops substantially for Russian text (0.8555 AUROC) compared to near-perfect scores for Urdu, Indonesian, and Arabic, suggesting under-training on Russian IC-338 Factuality enhancement methods (DoLa, ICD, ITI, TruthX, CD) cause large and consistent declines in context-faithfulness of LLaMA2-7B-Chat and LLaMA2-13B-Chat IC-340 GCG jailbreaking attacks exhibit strong model-specific transferability, achieving below 3% ASR on Llama-2-13b-chat and Llama-3.1-8b-instruct but above 90% ASR on Vicuna-13b-v1.5 and Mistral-7b-instruct IC-345 GPT-4o's table2latex outputs lose 63% of the time in human evaluation, with inconsistent formatting (lines, borders, margins) as the primary failure IC-348 Sequential parameter-modifying editing causes progressive degradation of general abilities in GPT-2 XL, Llama-2 7B, and Llama-3 8B, driven by growth in the condition number of the edited matrix IC-350 Editing conceptual knowledge with rome on Llama-2 7B is harder than factual knowledge editing, with the model failing to update concept-instance relationships IC-353 Llama-3.1-instruct 8b and 70b fail the harder NIAH test (sandwich needle) but pass the easier passkey retrieval test IC-362 LLMs with chain-of-thought prompting predict and simulate human risky choices that are more rational than actual human behavior, correlating more highly with maximum expected value than with human choices IC-367 Sparsifying initial tokens of the prefill phase causes disproportionate degradation in Llama-3-8B due to attention sink behavior IC-374 Pythia models are vulnerable to prompt injection via distractor text, and request-patching from a single trusted input restores most of their accuracy IC-385 TAR-bio-v1 retains bio-weaponization knowledge despite appearing to unlearn it; a different prompt template and answer extraction method reveals accuracy above 45% on WMDP-bio IC-393 GPT-4-turbo, Llama-3.1-8B-Instruct, and OpenAI Moderation show declining hate speech detection accuracy as sentence implicitness increases, with very low success rates in the highest implicitness ranges IC-397 Mistral 7B Instruct and Llama 3 8B Instruct exhibit systematic misalignment between their operational semantics of subjective phrases and human expectations, producing unexpected side effects when steered with certain phrases IC-402 In LLaMA3-8B, LLaMA2-13B, and Mistral-7B, soft-prompt information flow peaks in shallow layers (2–10) and reasoning correctness depends on whether deeper layers redirect attention away from soft prompts to earlier reasoning steps IC-407 Safety-aligned LLMs (Llama-2-chat, Llama-3-instruct, Gemma, GPT-3.5, GPT-4o, R2D2) achieve 100% jailbreak attack success rate under adaptive prompt-and-suffix attacks on 50 harmful requests IC-408 Claude models (2.0, 2.1, 3 Haiku, 3 Sonnet, 3 Opus, 3.5 Sonnet) achieve 100% jailbreak attack success rate under prefilling attacks via the Anthropic API IC-409 Knowledge editing methods correct verified hallucinations in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B far less effectively than their scores on existing benchmarks suggest IC-410 Knowledge editing can degrade generalization performance below pre-edit levels in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B IC-412 Edited knowledge in Llama2-7B is significantly less robust to adversarial prompts than in Llama3-8B and Mistral-v0.3-7B IC-414 LLaVA-1.5-7B and LLaVA-1.5-13B exhibit severe performance degradation when H2O KV cache compression is applied in multimodal settings IC-421 Sequential context-switching queries jailbreak Llama and Mistral models at 95% attack success rate IC-424 Structural in-context learning is transient in MultiBERTs and Pythia-1.4B, disappearing after early training IC-425 Pretrained GPT-2 Large fails at structural in-context learning on unseen tokens in a syllogism task IC-431 Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD dataset IC-433 LLMs exhibit a non-monotonic ID-OOD performance gap (generalization valley) that peaks at intermediate task complexity IC-436 API selection accuracy of 10 LLM-based agents degrades sharply as task complexity increases, with open-source models ≥70B matching closed-source on simpler tasks but lagging on the most complex IC-437 Extracting parameters from user queries is harder for LLM-based agents than using outputs from previous actions, and less intelligent LLMs show steeper parameter-filling degradation with task difficulty IC-438 All 10 LLM-based agents perform poorly at recognizing when they need to request input from the system or user, with overall accuracy between 30.55% and 55.18% IC-443 Model inconsistency on probing questions negatively correlates with skill-slice accuracy (r = -0.675), with models contradicting themselves more often on skills where they perform poorly IC-444 ViV1T produces non-differentiable population response representations when simulating mouse V1 experiments IC-445 Stable Diffusion v1.5 generates nudity for 796 out of 4703 prompts in the I2P inappropriate prompts dataset IC-448 CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasets IC-449 CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255 IC-456 VLM decoders achieve near-random accuracy on VALSE image-sentence alignment while pairwise accuracy is much higher, indicating reliance on linguistic priors IC-464 Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate method IC-466 GPT-4o's evaluation reliability degrades when used as the final judge in a collective thought pipeline that aggregates reviews from less reliable VLMs IC-467 Llama-3.1-405B's standard speculative decoding verification rejects correct continuations from GPT-4o, Llama-3.1-8B, and human text, accepting only roughly two tokens before the first rejection for GPT-4o IC-469 CLIP's global contrastive alignment causes attention on anatomically irrelevant regions in 3D CT, yielding limited zero-shot diagnostic accuracy (AUC 68.4 on 54 tasks) IC-470 LOVT and MGCA, which use implicit cross-attention local alignment, show only marginal improvement over CLIP in 3D CT diagnosis (AUC 69.4 and 70.1 vs 68.4) IC-474 GPT-4, GPT-4o, and Llama-3.1-405B fail at knowledge classification and comparison without chain-of-thought IC-475 GPT-4, GPT-3.5, GPT-4o, and Llama-3.1-405B fail at inverse knowledge search regardless of prompting IC-479 SAM's mask decoder exhibits attention drift to background or specific object parts under imprecise prompts, causing severe segmentation degradation IC-480 Pre-trained M-LLMs (GPT-4o, LLaVA-v1.6-34B, InternVL2-26B, Qwen2-VL-7B) produce imprecise tampering explanations when artifacts require fine-grained pixel-level analysis such as lighting or perspective inconsistencies IC-486 LLMs achieve correct final answers through incorrect reasoning chains, inflating their apparent multi-step reasoning performance IC-488 LLM performance degrades progressively as the number of reasoning hops increases, with error propagation from earlier sub-questions IC-489 State-of-the-art MLLMs fail at multi-step visual analogical reasoning, with best accuracy at 13% (Llama 3.2) on VOILA-WD and 29% (GPT-4o) on VOILA-ND, far below human performance of 71% and 70% IC-490 GPT-4o can identify visual relationships at 97% accuracy when given ground-truth descriptions but drops to 17% when asked to apply known relationships to new visuals, revealing a specific bottleneck in relational transfer IC-491 Presenting three images as a single collage rather than sequentially reduces MLLM accuracy by approximately 40% on the relationship application step IC-495 All evaluated multimodal foundation models achieve average non-hallucination accuracy below 50% across six hallucination scenarios IC-496 GPT-4o achieves the highest location inference accuracy among evaluated models, reaching 98.16% for country, 60.23% for city, and 27.13% for zip code from street view images IC-497 Text-to-image models experience performance drops exceeding 10% under adversarial prompts, with spatial reasoning being the most vulnerable task across all models IC-498 Multimodal foundation models exhibit severe group unfairness, with race and age biases more pronounced than gender bias in text-to-image models while gender bias is stronger in image-to-text models IC-508 GPT-4 exhibits reduced preference consistency (0.66 vs 0.84) when the quality distinction between two responses is minimal IC-510 Qwen2-7B and Llama3-8B score near-random on textual temporal reasoning tasks while Qwen2-72B, Llama3-70B, and GPT-4o achieve near-perfect accuracy, showing temporal reasoning in LLMs is scale-dependent and emerges only above ~70B parameters IC-520 Temporal reasoning performance drops significantly across all 15 MLLMs when video frames are shuffled or reduced to one-fifth of the original count IC-522 LLMs perform correct example inference without inducing the correct rule, and this gap is robust to prompting methods, fact count, and scenario form IC-528 The knowledge localization assumption fails for a large fraction of facts in GPT-2, Llama2-7B, and Llama3-8B, with 77% of facts classified as inconsistent knowledge in Llama3-8B IC-531 CONCH's zero-shot encoders cannot discriminate survival risk, achieving near-random concordance index on pathology whole-slide images IC-532 PLIP's zero-shot encoders produce random-guessing-level survival predictions and consistently underperform CONCH on pathology survival analysis IC-539 GPT-4o in text-code-image mode achieves the highest scores on SCIMAGE but remains below 4 on all three evaluation dimensions, and all models degrade substantially on prompts requiring combined understanding types IC-540 Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image models IC-542 The PyTorch pretrained ResNet50 on ImageNet is vulnerable to (1, y)-ACE calibration attacks that increase ECE from 3.70% to 47.23% while preserving accuracy IC-549 All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBench IC-551 GPT-4's workflow generation performance declines as the number of nodes and edges in the workflow increases IC-555 Large LLMs (GPT-3.5-turbo, Gemini 1.5 Flash, Llama3-70B, Mixtral 46.7B) exhibit reasoning errors and significant accuracy degradation on large-scale logical commonsense reasoning tasks with 32k+ rules, even when the knowledge base is complete and retrieval is ideal IC-557 Guard models show significantly degraded calibration under jailbreak attacks, with prompt classification ECE substantially higher than response classification ECE IC-558 Guard models exhibit inconsistent calibration when classifying responses from different response model types, with ECE varying by up to 39 percentage points within a single model IC-569 Mistral-7B-instruct-v0.1 achieves only F1 of 0.419 on zero-shot stance detection for the X-Stance German dataset, substantially below the fine-tuned BERT baseline (F1 0.693) IC-570 Gradient-based image jailbreaks optimized against single or ensemble VLMs are universal for the attacked model(s) but do not transfer to other VLMs, except between highly similar models IC-571 Open-source VLMs (LLaVA, MiniGPT-4, InstructBLIP) are substantially more vulnerable to multimodal jailbreak attacks than Gemini-1.5-flash, with BAP attack ASR of 58–62% versus 40–41% IC-572 Bijection learning achieves state-of-the-art jailbreak ASR on frontier models, with peak ASR increasing with model capability IC-573 Model capabilities on MMLU degrade monotonically as bijection encoding complexity increases IC-574 Guard models fail to effectively mitigate bijection attacks even at capability parity with the target model IC-586 Symbolic distance (number of reasoning steps) is the primary bottleneck for relational reasoning in LLMs, not total context length IC-589 Flavor text (non-essential descriptive language) degrades relational reasoning in most LLMs, but GPT-4o is robust to it IC-591 Downstream fine-tuning on GSM8K degrades safety of Llama2-7b-chat and Mistral-7b-instruct-v0.2, but RSN-Tune partially preserves safety by protecting non-overlapping safety neurons. IC-592 The log-likelihood layer in LLaMA-2-7B, LLaMA-2-7B-Chat, Vicuna-7B, and Mistral-7B-Instruct produces factually incorrect answers on TruthfulQA MC1 (817 samples) due to a misalignment between the output distribution and internal attention head representations, with LM-to-head-norm accuracy gaps of 24.23 to 40.68 points. IC-595 Gemini-pro has knowledge gaps on specific topics (Permian extinction, Fordism) causing it to perform far below its average rank on existing benchmarks IC-596 Multiple released LLMs fail to refuse harmful prompts disguised as historical or philosophical discussions, with GPT-4o and Mixtral showing the lowest refusal rates IC-598 GPT-4o's synthetic image detection accuracy drops sharply on specialized domains (satellite 45.0%, medical 54.3%) compared to common image types (object 84.4%, person 84.4%) IC-599 All evaluated audio LMMs perform at or near random chance (44.4%–51.2%) on synthetic audio detection, while humans achieve 69.2% IC-601 Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreases IC-603 SIREN's embedding layer is not effectively optimized by gradient descent, so its embedding frequencies must be set as a hyperparameter IC-605 Six SOTA LLMs are vulnerable to composable jailbreak attacks, with maximum attack success rates ranging from 44% to 94% IC-607 CLMBR-T-BASE's clinical prediction performance degrades as patient EHRs become more repetitive or irregular IC-610 In LLaMA2-7B-Chat, RAG hallucinations are causally driven by copying heads losing external context information during generation and by knowledge FFNs in mid-to-upper layers over-adding parametric knowledge to the residual stream IC-632 BLIP-2 succeeds on only 5 out of 100 advanced compositional vision-language tasks IC-637 KN edit (neuron suppression) has low reliability, overturning at most 5.2% of BLIMP categorical predictions and achieving only 1.66%–47.86% reliability on factual tasks IC-638 ROME editing on GPT-2 XL and Llama-2 7B achieves high reliability but fails under bijective symmetry (23.71%–33.64%) and synonymous invariance (52.35%–58.36%) criteria IC-647 DINOv2's feature-map artifacts cause it to be incompatible with the LOSt unsupervised object discovery method, scoring far below DINO IC-668 SAM's edge-oriented segmentation yields high recall but very low precision because it cannot distinguish object boundaries from interior edges IC-684 CLIP ViT-B/32 misclassifies 99% of forest satellite images as ocean when the word 'ocean' is overlaid as text IC-687 GPT-3.5+ models exhibit a gambler's fallacy bias and generate low-complexity sequences when asked to produce random binary sequences IC-698 GPT-3.5-turbo's CoT reasoning errors are correlated across different demonstration sets, while PoT errors are less correlated IC-699 Llama2-13b produces significantly less consistent answers than GPT-3.5-turbo on complex reasoning tasks, making it unsuitable as a weaker LLM in a cascade IC-700 GPT-4's reasoning accuracy degrades when provided with incorrect hints from a weaker model IC-713 Llama-2-7B-Instruct exhibits a reproducible failure mode in book-length summarization: high repetition and complete inability to perform incremental updating IC-716 Factual information deleted from GPT-J, LLaMA-2, and GPT-2-XL via ROME or MEMIT is recoverable by sampling outputs on automatically generated rephrased prompts, with up to 56% extraction success at budget b=20 IC-730 CLIPCap and BLIP-2 produce degraded alt-text on Twitter social media images, with BLEU@4 of 0.372 and 0.111 respectively IC-731 CLIP (ViT-B/32) achieves only 17.5 recall on video-text temporal alignment because it was trained on images and lacks video dynamics IC-733 CoT prompting improves factual accuracy for instruction-tuned LLMs but degrades it for non-instruction-tuned LLMs such as OPT, BLOOM, and LLaMA IC-734 GPT-3.5-turbo's factual verification F1 decreases as the number of reasoning hops required to validate a claim increases IC-735 GPT-3.5-turbo's factual verification performance drops substantially under adversarial modifications, with man-made adversarial examples causing the largest decline IC-737 CLIP ViT-B/16 binarized dot products yield 0.50–0.58 accuracy on binary concept presence queries across five image classification datasets IC-747 SOTA pruning methods (SparseGPT, Wanda, magnitude) cause significant degradation on knowledge-intensive tasks for Vicuna and Llama models at 25-30%+ unstructured sparsity, and fail completely for n:m structured sparsity IC-750 Open-source models without safety training are significantly more vulnerable to jailbreak attacks than safety-aligned proprietary models IC-761 GPT-3.5-turbo and GPT-4 are susceptible to specific circulating jailbreaking prompts, with 'jailmommy' achieving a 71.16% success rate in producing toxic outputs IC-762 GPT-4 and GPT-3.5 outperform humans in generation but underperform in discriminative (selective) evaluation across 10 of 13 language tasks IC-763 CLIP and OpenCLIP fall short of human discriminative accuracy on vision tasks, with performance dropping substantially under hard negatives IC-764 GPT-4 and GPT-3.5 make frequent errors answering questions about their own generated text, underperforming humans in interrogative evaluation IC-765 BLIP-2, BLIP, InstructBLIP, Bard, and BingChat fall short of human accuracy in answering questions about Midjourney-generated images IC-766 Multiple Real-SR methods fail to outperform a small FSRCNN network on the majority of 100 representative degradation cases IC-767 BSRNet outperforms RealESRNet on the majority of degradation cases, reversing the ranking obtained from a single random test set IC-783 Stable Diffusion v1.5's conditional probability pθ(x|c) is heavily biased by the unconditional probability pθ(x), making it unreliable as a condition-alignment metric IC-784 Pre-trained scoring models (CLIP Score, HPS, Image Reward, Pick Score) underperform on domain-specific fine-tuned diffusion models IC-788 CLIP ViT-B/32 fails to retrieve the correct image even when the generated target caption is well-aligned with the ground-truth image IC-807 Language-conditioned robot policies RT-1 and RT-2 fail to generalize to unseen manipulation tasks, achieving only 16.7% and 11.1% success rates respectively across 7 novel skills IC-812 InstructGPT (text-davinci-003) reduces content diversity in co-written essays while GPT-3 (davinci) does not, and the effect is attributable to the model's own less diverse text contributions IC-815 RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B models IC-816 RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B models IC-834 CLIP zero-shot predictions exhibit high equal opportunity difference when target and sensitive attributes are intrinsically dependent IC-840 GPT-2 Small's name mover heads exhibit disrupted attention patterns under out-of-distribution Gaussian noise corruption IC-848 GPT-3.5 achieves 98.3% precision and 96.0% recall for automatic question-tuple matching but makes errors when questions differ in wording yet are semantically unique IC-849 GPT-2 and T5-base exhibit vanishing expected gradients under RFT for inputs with small reward standard deviation, prevalent in 3 of 7 GRUE datasets, causing RFT to underperform SFT IC-852 Intrinsic self-correction without external feedback consistently degrades reasoning accuracy across GPT-3.5-turbo, GPT-4, GPT-4-turbo, and LLaMA-2-70B-chat IC-861 LLMs achieve near-zero accuracy on the disconnected nodes task, indicating an inability to reason about the absence of edges in a graph IC-871 Llama-2-70b-chat as a grader is more generous than GPT-4 and systematically gives higher scores to Llama-2 family outputs IC-887 HuggingGPT's in-context task-model assignment always selects the same model regardless of input question or task type IC-888 SLIMG fails on heterophily graphs for link prediction because it cannot properly measure node similarity of heterophily embeddings IC-898 GPT-3.5-turbo and GPT-4 are overconfident in their initial code predictions when unit test execution is unavailable IC-900 BLIP-2, MiniGPT-4, and LLaVA-1.5 show degraded zero-shot VQA accuracy on underspecified questions, with absolute improvements of 1.14–7.94% when questions are augmented with visually-grounded details IC-902 BLIP-2 and MiniGPT-4 confidence-based question selection underperforms the original question for paraphrased candidates but succeeds for semantically enriched REPARe questions IC-904 Stable Diffusion XL generates non-empty cups when prompted for 'empty cup' IC-927 Human ciphers (ASCII, Unicode, Caesar, Morse) bypass the safety alignment of GPT-4 and GPT-3.5-turbo, with more powerful models producing more unsafe responses IC-928 SelfCipher (a role-play prompt without explicit cipher rules) evokes a 'secret cipher' in LLMs, achieving high unsafety rates that outperform most human ciphers IC-930 StackLLama, when used as a reward model, achieves near-random consistency on contrast instructions for the Stack Exchange task IC-932 Pre-trained DNN object detectors show a sharp falloff in peripheral detection performance with increasing eccentricity, degrading to near-chance by 20°, while human performance degrades gradually IC-939 CLIP reward landscapes are well-shaped for photorealistic environments but poorly shaped for abstract renderings IC-940 CLIP can specify 5 of 8 complex humanoid tasks from single-sentence prompts, failing on tasks requiring discrimination of subtle body-pose differences IC-946 LLaMA-2-7B, MPT-7B, Falcon-7B, and Pythia-12B do not consistently improve in perplexity as the StreamingLLM cache size increases IC-981 ImageBind's indirect alignment through images degrades zero-shot performance on non-visual modalities and prevents emergent cross-modal retrieval IC-985 LLaMA-2-7B plateaus in ICL accuracy and fails to override semantic priors when in-context labels are flipped on a simple happy/sad classification task IC-986 Most LLMs lack tool usage awareness, with only ChatGPT exceeding 70% F1 in zero-shot evaluation IC-987 When the correct tool is absent from the candidate list, most LLMs hallucinate a tool rather than returning 'none' IC-988 LLMs show large gaps in multi-tool selection and over-rely on the number of tools specified in the prompt IC-989 Tool selection CSR degrades as the candidate tool list grows from 5 to 15 tools, and performance varies by user scenario IC-991 LLaMA-2, Falcon-7B, and GPT-3.5-Turbo exhibit large performance spread (up to 76 accuracy points) across semantically equivalent prompt formats, and model comparison rankings are frequently reversed by format choice IC-999 FLUX1 and Stable Diffusion 3.5 exhibit high local dependency ratio and produce text hallucinations when generating text content TM-003 Optical-to-SAR generation errors cluster spatially, but only five regions survive FDR control TM-004 Surface composition, not the acquisition time gap, correlates with optical-to-SAR reconstruction error