Comparing 27 CNNs and 14 transformers (all video action recognition models trained on Kinetics-400), the two architectures are roughly equivalent in maximum RSA score across brain regions. However, their alignment profiles across depth differ strikingly: in V2, transformers compute high-correlating representations at around 0.1 of total network depth, while CNNs reach comparable alignment only at around 0.6 of total depth. In EBA, CNNs show a gradual increase in correlation with depth, while transformers maintain relatively stable representations until the classification layer where the score increases abruptly. This pattern also holds for image object recognition CNNs and transformers.
Evidence
correlational
Key metric
In V2: transformers peak at ~0.1 of total depth, CNNs at ~0.6 of total depth; in EBA: CNNs show gradual increase, transformers stable until classification layer
Caveat
The paper notes this pattern was not observed in image fMRI studies (Nonaka et al. 2021) and hypothesizes it may relate to the dynamic nature of video stimuli; requires validation on other video fMRI datasets