IC-344GPT-4o achieves the highest average score (64.62) on general document benchmarks, outperforming Qwen2-VL-72B (58.40) and GeminiPro-1.5 (57.05)

Juan A. Rodriguez, Xiangru Jian, Siba Smarak Panigrahi, Tianyu Zhang, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Zichao Li, Suyuchen Wang, Pierre-Andre Noel, Mats Leon Richter, Saverio Vadacchino, Shubham Agarwal, Sanket Biswas, Sara Shanian, Ying Zhang, Sathwik Tejaswi Madhusudhan, Joao Monteiro, Krishnamurthy Dj Dvijotham, Torsten Scholak, Nicolas Chapados, Sepideh Kharaghani, Sean Hughes, M. Özsu, Siva Reddy, Marco Pedersoli, Yoshua Bengio, Christopher Pal, Issam H. Laradji, Spandana Gella, Perouz Taslakian, David Vazquez, Sai Rajeswar

SourceBigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks

On 12 established document understanding benchmarks (DocVQA, InfoVQA, DeepForm, KLC, WTQ, TabFact, ChartQA, TextVQA, MMMU, DUDEmini, SlideVQA-M, TableVQA), GPT-4o attains the top average score of 64.62. Qwen2-VL-72B follows at 58.40, GeminiPro-1.5 at 57.05, Claude-3.5 Sonnet at 50.73, and Llama-3.2-90B at 38.82. This contrasts with BigDocs-Bench where GPT-4o is outperformed by smaller open models.

Evidence
correlational
Key metric
GPT-4o avg 64.62, Qwen2-VL-72B 58.40, GeminiPro-1.5 57.05, Claude-3.5 Sonnet 50.73, Llama-3.2-90B 38.82
Caveat
Llama-3.2-90B scores use chain-of-thought prompting (marked with asterisk) for DocVQA and TabFact, which may inflate its score relative to other models evaluated zero-shot.
Model
GPT-4o, Qwen2-VL Qwen2-VL-72B, Gemini 1.5 / Gemini Pro 1.5, Claude 3.5 Sonnet, Llama 3.2 Llama-3.2-90B
Datasets
DocVQA [eval], ChartQA [eval], MMMU [eval]
Related findings
IC-343, IC-345
Extraction
automatic-extraction