IC-195Transformers achieve high brain alignment in early visual cortex at much shallower network depth than CNNs

Christina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris Groen

SourceOne Hundred Neural Networks and Brains Watching Videos: Lessons from Alignment

Comparing 27 CNNs and 14 transformers (all video action recognition models trained on Kinetics-400), the two architectures are roughly equivalent in maximum RSA score across brain regions. However, their alignment profiles across depth differ strikingly: in V2, transformers compute high-correlating representations at around 0.1 of total network depth, while CNNs reach comparable alignment only at around 0.6 of total depth. In EBA, CNNs show a gradual increase in correlation with depth, while transformers maintain relatively stable representations until the classification layer where the score increases abruptly. This pattern also holds for image object recognition CNNs and transformers.

Evidence
correlational
Key metric
In V2: transformers peak at ~0.1 of total depth, CNNs at ~0.6 of total depth; in EBA: CNNs show gradual increase, transformers stable until classification layer
Caveat
The paper notes this pattern was not observed in image fMRI studies (Nonaka et al. 2021) and hypothesizes it may relate to the dynamic nature of video stimuli; requires validation on other video fMRI datasets
Model
TSM, I3D, SlowFast, X3D, MViT V2, VideoMAE, Uniformer, TimesFormer
Concepts
Depth-dependent structure
Datasets
BOLD Moments Dataset [eval], Kinetics-400 [source]
Methods
Representational Similarity Analysis [primary]
Related findings
IC-194, IC-196, IC-197
Extraction
automatic-extraction