IC-1355Instruction-tuned VLMs fail to follow multiple-choice format in reasoning questions, with InstructBLIP frequently returning blank responses

Yuanfeng Ji, Chongjian GE, Weikai Kong, Enze Xie, Zhengying Liu, Zhenguo Li, Ping Luo

SourceLarge Language Models as Automated Aligners for benchmarking Vision-Language Models

On closed-ended (multiple-choice) reasoning questions, BLIP2 tends to select a particular option as instructed, while instruction-tuned models (InstructBLIP, LLaVA, LLaMA-Adapter v2, mPLUG-Owl, Otter) often ignore the provided options and generate detailed free-form responses. InstructBLIP in particular frequently returns blank responses. The authors attribute this to overfitting during instruction tuning, which causes catastrophic forgetting of generic instruction-following ability, making it difficult for the LLM judge to evaluate their outputs.

Evidence
observational
Caveat
This is an observational note from the authors during evaluation rather than a separately quantified experiment; no dedicated metric isolates format compliance from content accuracy.
Model
InstructBLIP, LLaVA, LLaMA-Adapter v2, mPLUG-Owl, Otter
Concepts
Failure mode
Datasets
Auto-Bench [eval]
Related findings
IC-1356
Extraction
automatic-extraction