IC-1509Kalamang-English translation performance on MTOb increases with model size within the Llama and Llama 2 families, and GPT-4 outperforms Text-davinci-003

Garrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, Luke Melas-Kyriazi

SourceA Benchmark for Learning to Translate a New Language from One Grammar Book

Across the Llama family, best CHRF on KGV→ENG rises from 16.1 (7B) to 13.2 (13B) to 29.7 (30B), and within Llama 2 from 26.8 (7B) to 30.4 (13B) to 37.8 (70B). The 13B Llama is an outlier that underperforms the 7B, but the overall trend is positive. Among API models, GPT-4 consistently matches or outperforms Text-davinci-003 across context settings (e.g., 38.4 vs 38.4 best KGV→ENG; 40.0 vs 34.8 best ENG→KGV). The paper notes that details of API-based models are not known, so the comparison is less controlled.

Evidence
correlational
Key metric
best chrf (kgv→eng): llama-7b 16.1, llama-13b 13.2, llama-30b 29.7; llama 2-7b 26.8, llama 2-13b 30.4, llama 2-70b 37.8; gpt-4 38.4 vs text-davinci-003 38.4 (kgv→eng), gpt-4 40.0 vs text-davinci-003 34.8 (eng→kgv)
Caveat
The Llama 13B underperforms the 7B, breaking the monotonic trend. API model training details are unknown, so cross-family comparisons are less controlled. Only one instruction-tuned variant per Llama 2 size was tested.
Model
LLaMA Llama 7B, LLaMA-13B, LLaMA-30B, Llama 2 / Llama 2 base Llama 2 7B, Llama 2 13B, Llama 2 70B, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3 / GPT base text-davinci-003
Concepts
Scale-dependent behaviour
Methods
CHRF [eval]
Related findings
IC-1508, IC-1510
Extraction
automatic-extraction