IC-1508LLMs with in-context learning translate Kalamang-English at 44.7/45.8 CHRF, falling short of the human baseline of 51.6/57.0 CHRF

Garrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, Luke Melas-Kyriazi

SourceA Benchmark for Learning to Translate a New Language from One Grammar Book

The paper evaluates 10 released LLMs on the MTOb benchmark, which requires translating between English and Kalamang using a 573-page grammar book and a bilingual word list as in-context reference materials. The best-performing model, Claude 2 with word list, parallel sentences, and a ~100k-token grammar book chunk in context, achieves 44.7 CHRF on KGV→ENG and 45.8 CHRF on ENG→KGV. A human who learned Kalamang from the same reference materials scores 51.6 and 57.0 CHRF respectively, a gap of +6.9 and +11.2 CHRF. Model outputs are described as erratic: some are near-perfect while others make nonsensical errors, in contrast to the human's consistently grammatical translations.

Evidence
correlational
Key metric
44.7 chrf (kgv→eng) and 45.8 chrf (eng→kgv) for claude 2 (w+s+gl), vs. 51.6 and 57.0 chrf human baseline
Caveat
The test set is drawn from the same field documentation as the training materials, so it may share vocabulary and structures with the reference book. Kalamang is a single language; generalisation to other typologically diverse languages is untested. The human baseline is a single individual with linguistics training.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, GPT-3 / GPT base text-davinci-003, Llama 2 / Llama 2 base Llama 2 70B, Llama 2 13B, Llama 2 7B, LLaMA LLaMA-30B, LLaMA-13B, Llama 7B
Datasets
A Grammar of Kalamang [source]
Methods
CHRF [eval], BLEU / BLEU@4 [eval]
Related work
BABYLM [context], FLORES-101 [context], XTREME-UP [context]
Related findings
IC-1509, IC-1510
Extraction
automatic-extraction