Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-582
Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIs
IC-583
Brain alignment in InstructBLIP and Idefics is organized by depth: middle layers align with higher visual regions while later layers align with early visual regions, whereas mPLUG-Owl shows later layers aligning with both
IC-584
Most brain-explained variance is shared across task instructions, with image captioning (IC) acting as an umbrella category showing high overlap with VQ and CR but lower overlap with iu2 and sr
IC-585
MLLMs effectively capture count-related and recognition-related visual concepts with distinct brain alignment patterns, but produce similar alignment patterns for color, positional understanding, and general scene understanding