Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
BLEU / BLEU@4
anchor
Findings
IC-1236
Among 7B LLMs, Llama-2-7b achieves the best zero-shot COMET scores in both translation directions, while MPT-7b leads in BLEU for en-to-xx
[eval]
IC-1237
Llama-2-7b's pre-existing translation knowledge is diluted by large amounts of parallel data, causing COMET to decline after 100k examples
[eval]
IC-1238
Llama-2-13b produces off-target non-translation outputs in zero-shot English-to-foreign-language translation
[eval]
IC-1508
LLMs with in-context learning translate Kalamang-English at 44.7/45.8 CHRF, falling short of the human baseline of 51.6/57.0 CHRF
[eval]
IC-304
Instruction-tuned LMs become more vulnerable to prompt-injected data extraction as model size increases from 7B to 70B
[eval]
IC-730
CLIPCap and BLIP-2 produce degraded alt-text on Twitter social media images, with BLEU@4 of 0.372 and 0.111 respectively
[eval]