IC-583Brain alignment in InstructBLIP and Idefics is organized by depth: middle layers align with higher visual regions while later layers align with early visual regions, whereas mPLUG-Owl shows later layers aligning with both
By mapping each cortical voxel to the layer (out of 33) that produces the highest normalized brain alignment, the paper reveals a depth-dependent structure in how MLLM representations correspond to visual processing. For InstructBLIP and Idefics, middle layers show greater alignment in higher visual regions (floc-bodies, floc-faces, floc-places, floc-words), while later layers align more with early visual regions (prf-visual ROIs). mPLUG-Owl differs: its later layers achieve higher alignment in both higher and early visual regions. The authors attribute this cross-model variation to differences in the underlying language decoder architectures.
Evidence
correlational
Caveat
This variation across the 3 MLLMs may be due to the difference in the underlying language decoder models, which generate output tokens, capture contextual representations, and influence the alignment trend across the layers