IC-662Translation ability in BLOOM models surges at approximately one-sixth of pre-training and then plateaus, with consistent dynamics across model sizes from 560M to 7.1B

Tingchen Fu, Lemao Liu, Deng Cai, Guoping Huang, Shuming Shi, Rui Yan

SourceThe Reasonableness Behind Unreasonable Translation Capability of Large Language Model

The paper examines officially released intermediate checkpoints of the BLOOM family (560M, 1.1B, 1.7B, 3B, 7.1B) and tracks bilingual translation perplexity across training steps on multiple language pairs (Chinese, Catalan, Eastern Panjabi, Igbo, Tswana). Translation ability experiences a surge at approximately 1/6 of the total training process, then plateaus or gradually increases. The pattern is consistent across all five model sizes, suggesting small models share the same underlying translation-learning mechanisms as large ones. Additionally, translation into low-resource languages is consistently worse than translation from low-resource languages into English, reflecting the importance of target-language modeling.

Evidence
observational
Key metric
surge at approximately 1/6 of the whole training process; warmup phase ends at approximately 1/1000 of the whole training process
Caveat
The surge cannot be solely attributed to learning rate changes since warmup ends at ~1/1000 of training; the paper does not provide a mechanistic explanation for why 1/6 specifically.
Model
BLOOM
Datasets
FLORES-200 [eval], WMT21 [eval]
Related work
BLOOM [builds-on]
Extraction
automatic-extraction