IC-582Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIs

SUBBA REDDY OOTA, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi GNVV, Manish Shrivastava, Maneesh Kumar Singh, Bapi Raju Surampudi, Manish Gupta

SourceCorrelating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

The paper trains bootstrap ridge regression encoding models to predict fMRI activity from model representations extracted via 10 natural instructions. Across the whole visual cortex and five functional localizers, all three MLLMs significantly outperform ViT-H (p≤0.05) and match or exceed CLIP-text, which uses ground-truth captions while MLLMs use predicted generations. In higher visual regions (floc-places), MLLM normalized alignment reaches approximately 0.8, dropping to approximately 0.6 in early visual ROIs. In Appendix K, non-instruction-tuned BLIP-2 performs marginally below CLIP-text, and LLaMA-2-7B (text-only, no visual input) performs closer to ViT-H, confirming that instruction tuning and multimodality drive the alignment gain.

Evidence
correlational
Key metric
ANOVA: p-value 0.008, f-statistic 14.60 (InstructBLIP, higher vs. early visual ROIs); p-value 0.009, f-statistic 13.85 (all MLLMs combined); normalized alignment ~0.8 in floc-places, ~0.6 in prf-visual ROIs
Caveat
CLIP-text uses golden oracle captions while MLLMs use mean pooling over predicted output tokens; the NSD dataset involves passive image viewing, which may not fully capture task-specific instruction processing
Model
InstructBLIP, mPLUG-Owl, Idefics, ViT ViT-H, CLIP / CLIP-ViT (LC), BLIP-2, Llama 2 / Llama 2 base Llama 2 7B
Datasets
NSD [eval], MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [source]
Methods
Principal component analysis [supporting]
Related findings
IC-583, IC-584, IC-585
Extraction
automatic-extraction