On TDIUC, which contains approximately 22% absurd questions unrelated to the image, all 10 LMMs show low zero-shot abstention F1 scores, indicating they tend to always give an answer. Increasing ICL shots substantially improves the ability to abstain: OFv2-9b absurd F1 rises from 29.02 (0-shot) to 56.44 (32-shot), and Idefics-9b from 26.87 to 67.45. However, even the best model (Idefics-9b (i)) still has relatively low F1. Instruction tuning and larger model size (up to 9B) further improve abstention.
The paper notes a positive correlation between overall accuracy and abstention performance, suggesting the two may not be fully independent. The evaluation only covers the case where the question is not relevant to the image, not other unanswerable scenarios.