IC-1080OPUS-MT small, OPUS-MT large, and mBART-50 in their default (unfine-tuned) form achieve low context-sensitive disambiguation accuracy on discourse-level phenomena

Gabriele Sarti, Grzegorz Chrupała, Malvina Nissim, Arianna Bisazza

SourceQuantifying the Plausibility of Context Reliance in Neural Machine Translation

The paper evaluates the released, unfine-tuned versions of OPUS-MT small, OPUS-MT large, and mBART-50 on English-to-French context-sensitive translation tasks covering anaphora resolution and lexical choice. Without context-aware fine-tuning, these models correctly disambiguate context-sensitive target words in only a minority of cases: OPUS-MT small scores 0.14/0.40/0.29, OPUS-MT large 0.16/0.41/0.31, and mBART-50 0.26/0.42/0.25 (ok metric on SCAT+, DisCEval-MT ana, DisCEval-MT lex respectively). This indicates that the released models have limited ability to leverage inter-sentential context for discourse-level disambiguation in translation.

Evidence
correlational
Key metric
ok: OPUS-MT small 0.14/0.40/0.29, OPUS-MT large 0.16/0.41/0.31, mBART-50 0.26/0.42/0.25 (SCAT+/DisCEval-MT ana/DisCEval-MT lex)
Caveat
Evaluation is limited to English-to-French and two discourse phenomena (anaphora resolution, lexical choice); the ok metric measures only whether the gold context-sensitive word appears in the output, not whether the model used context to arrive at it.
Model
OPUS-MT, mBART-50
Datasets
SCAT+ [eval], DisCEval-MT [eval]
Extraction
automatic-extraction