ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

2024-01-16 · ICLR 2024 poster · anchor

Findings