Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Multi-granularity Correspondence Learning from Long-term Noisy Videos
2024-01-16
· ICLR 2024 oral ·
anchor
Findings
IC-731
CLIP (ViT-B/32) achieves only 17.5 recall on video-text temporal alignment because it was trained on images and lacks video dynamics