Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
YouCook2
anchor
Findings
IC-048
Existing video LLMs (TimeChat, VTG-LLM, Momentor, Hawkeye) show limited zero-shot video temporal grounding capability and struggle to improve with fine-tuning
[eval]
IC-1389
Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMA
[source]