The paper evaluates zero-shot cross-lingual capabilities of XLM-R Large, XGLM-7.5b, BLOOM-7b, MPT-7b, Llama-2-7b, and Llama-2-13b on three benchmarks: XNLI (de, en, ru, zh), XStoryCloze (en, ru, zh), and XWinograd (en, ru, zh). Llama-2-13b achieves the highest averages on XStoryCloze (78.41) and XWinograd (79.05), while MPT-7b slightly leads on XNLI (46.58 vs 45.79). The paper states that Llama-2 demonstrates top performance for the tested languages overall.
Evaluations restricted to languages overlapping with fine-tuned languages (de, en, ru, zh); no dataset covers Icelandic; MPT-7b slightly outperforms Llama-2 on XNLI.