Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Something-Something V2
Findings
IC-1389
Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMA
[source]
IC-196
Models trained on Something-Something-V2 (which contains no faces) show reduced alignment with face-selective FFA compared to the same models trained on Kinetics-400
[source]