IC-194Temporal modeling in video models drives representational alignment to early visual cortex, while action classification task drives alignment to late brain areas
Christina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris Groen
Across 92 models (41 image object recognition, 10 image action recognition, 41 video action recognition all on Kinetics-400), video models score significantly higher than image models in early visual cortex (V1, V2, V3v, V3d), while in later areas the classification objective matters more: image models trained on action recognition outperform those trained on object recognition, and video models do not fare much better than action-trained image models. The control of image models trained on action recognition decouples temporal modeling from classification task, showing temporal modeling is the determining factor for early visual cortex alignment. In V2, the model MViT V2-S comes close to capturing almost all of the explainable variance for a portion of subjects.
Evidence
correlational
Key metric
Video models significantly outperform image models in early visual cortex; in late areas, action recognition models (both image and video) outperform object recognition models; MViT V2-S approaches upper noise ceiling in V2 for some subjects
Caveat
Results are based on a single fMRI dataset (BMD) and 10 subjects; the paper notes that a rival study (Garcia et al. 2024) found a different pattern with improved alignment mainly in MT and EBA, possibly due to methodological differences