IC-1239Llama-2-7b and Llama-2-13b achieve top zero-shot cross-lingual performance among 7B models on XNLI, XStoryCloze, and XWinograd

Haoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan Awadalla

SourceA Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

The paper evaluates zero-shot cross-lingual capabilities of XLM-R Large, XGLM-7.5b, BLOOM-7b, MPT-7b, Llama-2-7b, and Llama-2-13b on three benchmarks: XNLI (de, en, ru, zh), XStoryCloze (en, ru, zh), and XWinograd (en, ru, zh). Llama-2-13b achieves the highest averages on XStoryCloze (78.41) and XWinograd (79.05), while MPT-7b slightly leads on XNLI (46.58 vs 45.79). The paper states that Llama-2 demonstrates top performance for the tested languages overall.

Evidence
correlational
Key metric
Llama-2-13b: XNLI avg 45.79, XStoryCloze avg 78.41, XWinograd avg 79.05; MPT-7b: XNLI avg 46.58, XStoryCloze avg 76.23, XWinograd avg 76.23; XLM-R Large: XNLI avg 31.10, XStoryCloze avg 48.18, XWinograd avg 46.75
Caveat
Evaluations restricted to languages overlapping with fine-tuned languages (de, en, ru, zh); no dataset covers Icelandic; MPT-7b slightly outperforms Llama-2 on XNLI.
Model
Llama 2 / Llama 2 base Llama 2 7B, Llama 2 13B, Xlm-R XLM-R Large, XGLM XGLM-7.5B, BLOOM BLOOM-7B, MPT MPT-7B
Datasets
XNLI [eval], XWinograd [eval]
Methods
LM Evaluation Harness [eval]
Related findings
IC-1236, IC-1237, IC-1238
Extraction
automatic-extraction