IC-583Brain alignment in InstructBLIP and Idefics is organized by depth: middle layers align with higher visual regions while later layers align with early visual regions, whereas mPLUG-Owl shows later layers aligning with both

SUBBA REDDY OOTA, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi GNVV, Manish Shrivastava, Maneesh Kumar Singh, Bapi Raju Surampudi, Manish Gupta

SourceCorrelating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

By mapping each cortical voxel to the layer (out of 33) that produces the highest normalized brain alignment, the paper reveals a depth-dependent structure in how MLLM representations correspond to visual processing. For InstructBLIP and Idefics, middle layers show greater alignment in higher visual regions (floc-bodies, floc-faces, floc-places, floc-words), while later layers align more with early visual regions (prf-visual ROIs). mPLUG-Owl differs: its later layers achieve higher alignment in both higher and early visual regions. The authors attribute this cross-model variation to differences in the underlying language decoder architectures.

Evidence
correlational
Caveat
This variation across the 3 MLLMs may be due to the difference in the underlying language decoder models, which generate output tokens, capture contextual representations, and influence the alignment trend across the layers
Model
InstructBLIP, mPLUG-Owl, Idefics
Concepts
Depth-dependent structure
Datasets
NSD [eval]
Related findings
IC-582, IC-584, IC-585
Extraction
automatic-extraction