IC-1236Among 7B LLMs, Llama-2-7b achieves the best zero-shot COMET scores in both translation directions, while MPT-7b leads in BLEU for en-to-xx

Haoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan Awadalla

SourceA Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

The paper evaluates zero-shot translation performance of six 7B LLMs (OPT-7b, BLOOM-7b, Falcon-7b, Llama-1-7b, MPT-7b, Llama-2-7b) plus GPT-3.5-d and GPT-3.5-t on 10 translation directions (cs, de, is, zh, ru to/from English) using WMT'21 and WMT'22 test sets. Llama-2-7b achieves the highest average COMET in both directions (77.16 xx-to-en, 68.88 en-to-xx), while MPT-7b leads in BLEU for en-to-xx (15.90). All 7B models trail GPT-3.5-d and GPT-3.5-t by a substantial margin.

Evidence
correlational
Key metric
Llama-2-7b avg COMET 77.16 (xx-en) / 68.88 (en-xx); MPT-7b avg BLEU 23.90 (xx-en) / 15.90 (en-xx); GPT-3.5-d avg COMET 83.90 / 84.59; GPT-3.5-t avg COMET 85.46 / 86.56
Caveat
Icelandic test data is from WMT'21 while others are from WMT'22; beam size 5; the paper notes it relies more on COMET than BLEU due to better alignment with human evaluations.
Model
Llama 2 / Llama 2 base Llama 2 7B, MPT MPT-7B, OPT OPT-7B, BLOOM BLOOM-7B, Falcon Falcon-7B, Llama-1-7B, GPT-3.5 / ChatGPT-3.5 GPT-3.5-text-davinci-003, GPT-3.5-turbo-0301
Datasets
WMT21 [eval]
Methods
COMET [eval], BLEU / BLEU@4 [eval]
Related findings
IC-1237, IC-1238, IC-1239
Extraction
automatic-extraction