Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-632
BLIP-2 succeeds on only 5 out of 100 advanced compositional vision-language tasks
IC-633
BLIP-2, LLaVA, and mPLUG-Owl show a trade-off between caption length and hallucination rate on COCO
IC-634
Video-ChatGPT achieves the highest correctness, detail, and contextual scores among four video understanding baselines
IC-635
InstructBLIP achieves the highest overall MMBench score (44.0) among five evaluated vision-language models