Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
XNLI
anchor
Findings
IC-1054
Code fine-tuning degrades English natural language reasoning in Code LLaMA relative to LLaMA-2, but the effect is negligible or slightly positive in French, Spanish, and German
[eval]
IC-1239
Llama-2-7b and Llama-2-13b achieve top zero-shot cross-lingual performance among 7B models on XNLI, XStoryCloze, and XWinograd
[eval]
IC-1578
XLM-R-XL without instruction tuning produces [pad] tokens and fails to complete instruction-following tasks
[eval]