IC-1390Proficiency tests reveal that a substantial portion of correct main-test predictions by VidLMs and ILMs are spurious rather than reflecting robust understanding
Ilker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna, Emre Can Acikgoz, Letitia Parcalabescu, Iacer Calixto, Anette Frank, Albert Gatt, Aykut Erdem, Erkut Erdem
For each main test, VILMA includes a simpler proficiency test that checks a prerequisite capability (e.g., object identification, action recognition). The combined P+T metric counts a main-test success only if the proficiency test is also passed. A striking performance drop occurs from T to P+T across models and tasks, indicating that many apparently correct main-test predictions are made by chance or through reliance on spurious features. For example, VideoCLIP scores 50.8 on the change-of-state main test but only 25.9 on the combined P+T, and UniVL drops from 43.6 to 32.2 on action counting.
The proficiency tests are simpler than the main tests and may not perfectly capture the prerequisite capability. The paper acknowledges that the combined metric is stricter and that some drop may reflect the difficulty gap between proficiency and main tests rather than purely spurious reasoning.