Light Dark Scale-dependent behaviour A property holds in some size variants of a model family and not in others. Findings of this kind are statements about a variant rather than about an architecture, and they are only interpretable when the variant is recorded.
Findings IC-030 Larger language models exhibit higher independence rates and lower conformity under some protocols IC-058 The answer to factual questions appears in the top 0.01% most influential pretraining documents for 55% of 7B queries and 30% of 35B queries, but almost never for reasoning questions IC-062 Llama 3 70B implements temporal difference learning in-context for reward-based RL, with causally relevant SAE features in its residual stream, while Llama 3 8B performs at chance IC-072 LLM agents of varying scales exhibit a failure mode on web automation tasks when processing raw, complex web page observations, with the penalty being more severe for smaller models IC-097 Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3B IC-1008 The magnitude of CRITIC's improvement on mathematical program synthesis scales with Llama-2 model size IC-110 Phi-3 (3.8B) and Llama-3 (8B) with triplet prompting outperform GPT-4 with pairwise prompting on causal graph orientation IC-1136 Larger LLMs (LLaMA2-70B, Vicuna-33B) are more stubborn than their smaller counterparts (LLaMA2-7B, Vicuna-7B) when encountering incoherent entity-substitution counter-memory IC-1147 GPT-4 outperforms GPT-3.5 on all KITAB metrics but the gap is modest, with all-correctness below 35% for both, suggesting scale alone does not resolve constraint satisfaction IC-115 NUPA performance is largely independent of model size within a family: GPT-4o and GPT-4o-mini show nearly identical performance, as do Qwen2-72B and Qwen2-7B IC-1188 GPT-3 procedural planning performance is scale-dependent, with Curie (6.7B) scoring 3.75 and Davinci (175B) scoring 4.90 overall quality in few-shot settings, while GPT-4 achieves 4.81 overall and 5.00 order IC-119 The number of salient hallucination heads decreases as model size increases within the LLaVA family IC-1208 Larger LLMs (Llama-2-13B) require more data samples for successful backdoor injection via parameter editing compared to smaller models (GPT-2-XL 1.5B) IC-121 Gemini 1.0 Pro, when prompted as a zero-shot chain-of-thought judge with majority voting, underperforms fine-tuned smaller Gemma models as verifiers on GSM8K IC-1221 LLM performance on in-context boolean function learning is scale-dependent, with GPT-2 failing and Llama-2 models improving gradually with size IC-1252 Circuit overlap between IOI and colored objects in GPT-2 decreases as model scale increases from medium to xl IC-1261 LLaMA-2 attention to constraint tokens correlates with factual correctness, and a linear probe on these attention weights predicts factual errors comparably to model confidence IC-1262 LLaMA-2 factual query accuracy improves with entity popularity and decreases with query constrainedness, with larger models showing better performance on less popular and more constrained queries IC-1263 LLaMA-2 7B and 13B attention signal for predicting factual errors is available by approximately 50% of layers, enabling early stopping without performance degradation, while LLaMA-2 70B shows a slight performance drop IC-1265 Calibration and failure prediction improve as model capability scales from GPT-3 to GPT-4, but remain far from ideal IC-129 Sycophancy in VLMs increases with model size: InternVL-1.5-26B (95.8/89.6/86.5%) is more sycophantic than InternVL-1.5-2B (75.6/66.8/98.1%), and InternLM-XComposer2-VL-7B (36.7/28.0/50.7%) more than the 1.8B variant (33.3/20.2/33.0%) IC-1310 Chain-of-thought prompting provides notable code-optimization gains only for larger models (CodeLlama 34B, GPT-3.5, GPT-4) but not for CodeLlama 7B or 13B, consistent with an emergent capability IC-1317 Llama-2 and Pythia models contain linear representations of space and time that improve with depth and model scale IC-1319 Larger LLaMA and LLaMA2 models show better calibration on phrase-level tasks but not consistently on sentence- and paragraph-level tasks IC-1320 GPT-2 XL (1.5B) exhibits better calibration than larger models from the LLaMA, LLaMA2, and GPT-J families despite having fewer parameters IC-135 Linear relational embeddings for factual relations form in OLMo-7B, OLMo-1B, and GPT-J when subject-object co-occurrence frequency exceeds model-specific thresholds, with r=0.82 correlation between log co-occurrence and causality across all pretraining stages IC-1386 Fact recall in OPT and LLaMA models degrades by more than 5% relative accuracy when more than 30% of weights are pruned, and similarly when moving from the 30B to the 13B dense model IC-1400 When steerable feature dimension is held constant, increasing the type-l of steerable features does not improve performance of ESCN or EquiformerV2 on IS2RE and S2EF molecular property prediction IC-1408 OpenFlamingo and Idefics models generate low-quality explanations in zero-shot, but ICL and model scale significantly improve explanation CIDEr IC-1431 BERT-base fine-tuning has negligible distribution-wise variance (0.21%) while BERT-large fine-tuning has substantial distribution-wise variance (2.08%) on MRPC IC-146 Pythia models show increasing robustness to off-policy RLHF data as policy size scales from 410M to 2.8B IC-1498 The benefit of frozen LLM transformer blocks for visual encoding is scale-dependent: OPT blocks below 1.3B parameters degrade ViT-s performance while blocks at 1.3B and above improve it IC-1499 Self-rationalization quality and task accuracy scale with model size across GPT-3, FLAN-T5, and LLaMA on five QA datasets IC-1509 Kalamang-English translation performance on MTOb increases with model size within the Llama and Llama 2 families, and GPT-4 outperforms Text-davinci-003 IC-1598 Retrieval augmentation improves GPT-3.5-turbo-4k on long-context tasks but not GPT-3.5-turbo-16k IC-1631 Binding id mechanism fidelity increases with model size in both Llama and Pythia families IC-175 Context-sufficiency performance is scale-dependent: larger LLMs achieve high accuracy with sufficient context but still answer correctly 35-62% of the time without it, while smaller models hallucinate or abstain even with sufficient context IC-178 Self-guided confidence reasoning (SCR) outperforms rule-based confidence reasoning (RCR) for GPT-4o and GPT-4o mini, but RCR outperforms SCR for Llama-3-8B IC-197 Model computational complexity (FLOPs) shows a significant negative correlation with brain alignment in high-level brain areas IC-243 Bias ratio on certain multi-turn fairness tasks decreases with model size in the Gemma-2 and Qwen2.5 families IC-258 DNN accuracy on 3D perception tasks correlates with ImageNet object classification accuracy, suggesting 3D cues emerge as a byproduct of object recognition training IC-277 Open-source VLMs show a clear scaling trend in both average accuracy and reasoning robustness on DynaMath, with larger models performing substantially better IC-282 GPT-2 XL (1.5B) exhibits lower accuracy but reduced overconfidence (smaller ECE and Brier scores) compared to larger models on the CAT benchmark IC-296 The degree to which SAE features are active at multiple residual-stream layers increases with model size in Pythia, Gemma 2, Llama 3.2, and GPT-2 IC-302 Llama-3.1 models perform between unigram-inference and bigram-inference on Markov chain ICL, with performance improving monotonically with model scale IC-304 Instruction-tuned LMs become more vulnerable to prompt-injected data extraction as model size increases from 7B to 70B IC-315 CoT prompting (reasoning + instruction) yields larger relative gains for larger LLMs and harder problems in competitive code generation, with the effect reversing for the most capable models IC-319 Different CLIP backbones exhibit distinct robustness profiles to specific image perturbations, with each architecture being most resilient to a different transformation IC-320 Llama-2-13b-chat underperforms Llama-2-7b-chat on fine-grained dimension-level evaluation IC-326 Jailbreak vulnerability and response style are scale-dependent: Llama-2-13b-chat shows higher post-attack JSR and a shift toward direct compliance (type-5 actions) compared to Llama-2-7b-chat IC-335 GPT-4's detection performance as a scoring model is highly sensitive to the prompt, varying from 0.7289 to 0.9682 AUROC, far more than GPT-3.5 or Babbage IC-336 Larger proprietary LLMs (GPT-3.5 175B) are more effective universal text detectors than smaller models (Babbage 1.3B, GPT-Neo-2.7B), contradicting prior findings that smaller models are better IC-349 Larger LLMs (Llama-2 7B, Llama-3 8B) suffer more severe general ability degradation than smaller models (GPT-2 XL 1.5B) under the same number of sequential edits IC-365 70B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsity IC-368 Larger LMs (GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro) exhibit better calibration than their smaller counterparts (GPT-4o-mini, Claude-3-Haiku, Gemini-1.5-Flash) when verbalizing confidence with certainty phrases IC-386 LLM performance on CS-Bench grows logarithmically with parameter scale within model families IC-389 All evaluated LLMs score significantly lower on CS reasoning questions than knowledge questions, with the gap narrowing for stronger models IC-391 Llama-3 and Qwen-1.5 models exhibit position bias in LM-as-a-judge, retrieval-augmented QA, and math reasoning, with larger models showing less bias IC-396 Closed-source LLMs (GPT, Claude) rely more on deep structure than open-source LLMs (Llama, Mistral), and open-source models' surface sensitivity decreases with model scale IC-412 Edited knowledge in Llama2-7B is significantly less robust to adversarial prompts than in Llama3-8B and Mistral-v0.3-7B IC-422 Safety fine-tuning in Llama models improves with parameter size but exhibits diminishing returns IC-429 The point of emergence for MMLU shifts from approximately 1.3×10²² flops to 5.6×10²⁰ flops as models train on 64,000 task-relevant examples, and the log-linear fit R² improves from 0.632 to 0.950 IC-432 56 LLMs from 19 families exhibit u-shaped scaling on hard questions and inverted-U scaling on easy questions, with the opposing trends explaining emergent ability stagnation IC-434 The critical complexity at which LLMs over-rely on memorization shifts to higher task difficulty as model size increases IC-441 Skill-level improvements between model releases are highly uneven, with Claude 3.5 Sonnet gaining ~50% over Claude 3 Opus on law skills while Gemini improved most in math and science IC-455 Qwen-Audio 7B zero-shot underperforms a 128M parameter baseline on audio difference explanation across three evaluation scenarios IC-459 Qwen-2.5 models show that the privacy-utility tradeoff for differentially private steering improves with model size IC-471 Pythia models exceeding 100M parameters show a consistent leftward shift of the multifractal spectrum (increasing regularity) during training that is absent in the 14M and 31M variants IC-473 ResNet-18 lacks a clear multifractal structure while ResNet-152 shows one with irregular shifts, and a 160M diffusion model exhibits lower degree of emergence than Pythia 160M IC-510 Qwen2-7B and Llama3-8B score near-random on textual temporal reasoning tasks while Qwen2-72B, Llama3-70B, and GPT-4o achieve near-perfect accuracy, showing temporal reasoning in LLMs is scale-dependent and emerges only above ~70B parameters IC-537 Moirai's architectural enhancements (any-variate attention, multi-scale patch embedding, diverse mixture distribution) improve in-distribution forecasting but reduce out-of-distribution scalability relative to a simpler encoder-only baseline IC-550 Workflow generation performance scales with model size within families, but recently released 7B models outperform older 13B models IC-572 Bijection learning achieves state-of-the-art jailbreak ASR on frontier models, with peak ASR increasing with model capability IC-601 Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreases IC-606 The relationship between model size and jailbreak vulnerability is reversed between Anthropic and Meta model families IC-646 DINOv2, DeiT-III, and OpenCLIP repurpose approximately 2% of patch tokens in low-informative background areas as internal registers, discarding local patch information while aggregating global image information; DINO does not exhibit this behaviour IC-689 Subjective randomness generation and sharp ICL transitions emerge only in larger or reward-fine-tuned models, absent in earlier GPT-3 variants and smaller open-source models IC-717 Llama-2-Chat's evaluation capability does not improve monotonically with model size IC-736 Vicuna-13B outperforms Vicuna-7B on factual knowledge tasks by 5.4% on average IC-749 Compressed Vicuna-13B at 46.16% sparsity (matching 7B parameter count) achieves lower MMLU accuracy than dense Vicuna-7B, indicating large-sparse models do not outperform small-dense at matched size IC-789 CLIP retrieval quality in zero-shot compositional image retrieval scales log-linearly with model size from approximately 150M to 2.5B parameters IC-827 LLMs exhibit distinct psychological profiles that differ from human norms and vary by model size and version IC-837 CLIP ViT-L/14 embeds more target-attribute information and less sensitive-attribute information than CLIP ResNet-50 on Waterbirds IC-860 Larger PaLM 2 models (xxs to l) show progressively better graph reasoning, but even the largest variant fails to beat the majority baseline on edge existence IC-868 Skill-Mix performance degrades with increasing k, and within the Llama-2 family the saturation point increases with model size IC-903 Knowledge editing performance (ES, GS, LS) improves as model scale increases from GPT-2 (124M) to T5-XL (2.8B) to GPT-J (6B) across all editing methods IC-916 Pythia models show scale-dependent last-layer averaging barriers: 70M exhibits a barrier of ~13 while 410M shows ~1 IC-941 CLIP reward model quality scales with model size, with a sharp phase transition between ViT-H/14 and ViT-BigG/14 for the humanoid kneeling task IC-959 LLaMA and OPT-1.3B (and Aquila-7B) encode more similar interaction primitives than smaller models such as BERT-base and BERT-large TM-007 Coordinates are linearly decodable only in the larger TerraMind variants