IC-1406OpenFlamingo and Idefics models rarely abstain from answering unanswerable questions, but ICL significantly improves abstention F1

Mustafa Shukor, Alexandre Rame, Corentin Dancette, Matthieu Cord

SourceBeyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning

On TDIUC, which contains approximately 22% absurd questions unrelated to the image, all 10 LMMs show low zero-shot abstention F1 scores, indicating they tend to always give an answer. Increasing ICL shots substantially improves the ability to abstain: OFv2-9b absurd F1 rises from 29.02 (0-shot) to 56.44 (32-shot), and Idefics-9b from 26.87 to 67.45. However, even the best model (Idefics-9b (i)) still has relatively low F1. Instruction tuning and larger model size (up to 9B) further improve abstention.

Evidence
correlational
Key metric
OFv2-9b absurd F1: 29.02 (0-shot), 28.27 (4-shot), 42.02 (8-shot), 51.80 (16-shot), 56.44 (32-shot); Idefics-9b absurd F1: 26.87 (0-shot), 32.00 (4-shot), 47.51 (8-shot), 60.22 (16-shot), 67.45 (32-shot)
Caveat
The paper notes a positive correlation between overall accuracy and abstention performance, suggesting the two may not be fully independent. The evaluation only covers the case where the question is not relevant to the image, not other unanswerable scenarios.
Model
OpenFlamingo, Idefics
Concepts
Failure mode
Datasets
TDIUC [eval]
Methods
In-Context Learning / In-context learning prompt [primary]
Related work
OpenFlamingo [context], Idefics [context]
Related findings
IC-1405, IC-1407, IC-1408
Extraction
automatic-extraction