On 12 established document understanding benchmarks (DocVQA, InfoVQA, DeepForm, KLC, WTQ, TabFact, ChartQA, TextVQA, MMMU, DUDEmini, SlideVQA-M, TableVQA), GPT-4o attains the top average score of 64.62. Qwen2-VL-72B follows at 58.40, GeminiPro-1.5 at 57.05, Claude-3.5 Sonnet at 50.73, and Llama-3.2-90B at 38.82. This contrasts with BigDocs-Bench where GPT-4o is outperformed by smaller open models.
Llama-3.2-90B scores use chain-of-thought prompting (marked with asterisk) for DocVQA and TabFact, which may inflate its score relative to other models evaluated zero-shot.