IC-634Video-ChatGPT achieves the highest correctness, detail, and contextual scores among four video understanding baselines

Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny

SourceMiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

The paper evaluates four released video understanding models on the Video-based Generative Performance Benchmark. Video-ChatGPT scores highest on correctness (2.40), detail (2.52), contextual (2.62), and consistency (2.37) among the baselines. VideoChat scores 2.23/2.50/2.53/1.94/2.24, LLaMA Adapter 2.03/2.32/2.30/1.98/2.15, and Video-LLaMA 1.96/2.18/2.16/1.82/1.79. The paper identifies Video-ChatGPT as the strongest baseline.

Evidence
correlational
Key metric
Video-ChatGPT: correctness 2.40, detail 2.52, contextual 2.62, temporal 1.98, consistency 2.37; VideoChat: 2.23, 2.50, 2.53, 1.94, 2.24; LLaMA Adapter: 2.03, 2.32, 2.30, 1.98, 2.15; Video-LLaMA: 1.96, 2.18, 2.16, 1.82, 1.79
Model
Video-ChatGPT, VideoChat, LLaMA-Adapter, Video-LLaMA
Datasets
VideoInstruct100k [eval]
Related findings
IC-632, IC-633, IC-635
Extraction
automatic-extraction