Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
LM Evaluation Harness
anchor
Findings
IC-1054
Code fine-tuning degrades English natural language reasoning in Code LLaMA relative to LLaMA-2, but the effect is negligible or slightly positive in French, Spanish, and German
[eval]
IC-1239
Llama-2-7b and Llama-2-13b achieve top zero-shot cross-lingual performance among 7B models on XNLI, XStoryCloze, and XWinograd
[eval]
IC-427
Newer base models (post-November 2023) outperform older ones by 7.3 points on MMLU and 19.1 points on GSM8K controlling for pretraining compute, but this gap vanishes after fine-tuning all models on the same task-relevant data
[eval]
IC-432
56 LLMs from 19 families exhibit u-shaped scaling on hard questions and inverted-U scaling on easy questions, with the opposing trends explaining emergent ability stagnation
[eval]