SourceThe Reasonableness Behind Unreasonable Translation Capability of Large Language Model
The paper examines officially released intermediate checkpoints of the BLOOM family (560M, 1.1B, 1.7B, 3B, 7.1B) and tracks bilingual translation perplexity across training steps on multiple language pairs (Chinese, Catalan, Eastern Panjabi, Igbo, Tswana). Translation ability experiences a surge at approximately 1/6 of the total training process, then plateaus or gradually increases. The pattern is consistent across all five model sizes, suggesting small models share the same underlying translation-learning mechanisms as large ones. Additionally, translation into low-resource languages is consistently worse than translation from low-resource languages into English, reflecting the importance of target-language modeling.