SourceCorrelating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
Grouping the 10 instructions into five visual concept categories, the paper finds that count-related instructions (iu2, iu3, vq2) and recognition-related instructions (vq1, vq2) produce distinct brain alignment patterns: vq2 shows higher alignment in high-level visual regions for counting, while iu2 and iu3 show higher alignment in early visual regions. Recognition instructions show distributed alignment across both high-level and early regions. In contrast, instructions for color (iu1), positional understanding (vq3), and general scene understanding (cr) produce similar brain alignment patterns regardless of the specific concept, indicating MLLMs do not differentiate these visual concepts in their representations.