SourceLarge Language Models as Automated Aligners for benchmarking Vision-Language Models
On closed-ended (multiple-choice) reasoning questions, BLIP2 tends to select a particular option as instructed, while instruction-tuned models (InstructBLIP, LLaVA, LLaMA-Adapter v2, mPLUG-Owl, Otter) often ignore the provided options and generate detailed free-form responses. InstructBLIP in particular frequently returns blank responses. The authors attribute this to overfitting during instruction tuning, which causes catastrophic forgetting of generic instruction-following ability, making it difficult for the LLM judge to evaluate their outputs.