anchor
Findings
- IC-582Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIs [eval]
- IC-583Brain alignment in InstructBLIP and Idefics is organized by depth: middle layers align with higher visual regions while later layers align with early visual regions, whereas mPLUG-Owl shows later layers aligning with both [eval]
- IC-584Most brain-explained variance is shared across task instructions, with image captioning (IC) acting as an umbrella category showing high overlap with VQ and CR but lower overlap with iu2 and sr [eval]
- IC-585MLLMs effectively capture count-related and recognition-related visual concepts with distinct brain alignment patterns, but produce similar alignment patterns for color, positional understanding, and general scene understanding [eval]
- IC-666DINOv2, CLIP-vision, and VGG-19 representations all align with MEG brain responses, with DINOv2 showing particularly high retrieval performance for late brain activity after image offset [eval]