IC-1578XLM-R-XL without instruction tuning produces [pad] tokens and fails to complete instruction-following tasks

yisheng xiao, Juntao Li, Zechen Sun, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang, Min Zhang

SourceAre Bert Family Good Instruction Followers? A Study on Their Potential And Limitations

The paper evaluates the raw XLM-R-XL model (no instruction fine-tuning) on xcopa, xnli, xwinograd, and WMT'14 machine translation. On all three classification tasks the model outputs the special [pad] token instead of a meaningful answer. On machine translation it generates text unrelated to the prompt (e.g., 'fragen und antworten zum.' for a German-to-English translation request). The authors state that it is 'difficult or even impossible to collect the evaluation results' due to this pervasive failure, and provide only a few representative examples rather than a full quantitative evaluation.

Evidence
observational
Caveat
Only a few representative examples are shown (Table 19) rather than a systematic evaluation over the full test set; the authors note it is 'difficult or even impossible to collect the evaluation results' quantitatively.
Model
Xlm-R XLM-R-XL
Concepts
Failure mode
Datasets
XCOPA [eval], XNLI [eval], XWinograd [eval]
Extraction
automatic-extraction