SourceAre Bert Family Good Instruction Followers? A Study on Their Potential And Limitations
The paper evaluates the raw XLM-R-XL model (no instruction fine-tuning) on xcopa, xnli, xwinograd, and WMT'14 machine translation. On all three classification tasks the model outputs the special [pad] token instead of a meaningful answer. On machine translation it generates text unrelated to the prompt (e.g., 'fragen und antworten zum.' for a German-to-English translation request). The authors state that it is 'difficult or even impossible to collect the evaluation results' due to this pervasive failure, and provide only a few representative examples rather than a full quantitative evaluation.