IC-343GPT-4o, Claude-3.5 Sonnet, and GeminiPro-1.5 score below BigDocs-trained open models on BigDocs-Bench tasks requiring long structured code generation

Juan A. Rodriguez, Xiangru Jian, Siba Smarak Panigrahi, Tianyu Zhang, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Zichao Li, Suyuchen Wang, Pierre-Andre Noel, Mats Leon Richter, Saverio Vadacchino, Shubham Agarwal, Sanket Biswas, Sara Shanian, Ying Zhang, Sathwik Tejaswi Madhusudhan, Joao Monteiro, Krishnamurthy Dj Dvijotham, Torsten Scholak, Nicolas Chapados, Sepideh Kharaghani, Sean Hughes, M. Özsu, Siva Reddy, Marco Pedersoli, Yoshua Bengio, Christopher Pal, Issam H. Laradji, Spandana Gella, Perouz Taslakian, David Vazquez, Sai Rajeswar

SourceBigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks

The paper evaluates released models on BigDocs-Bench, a suite of 10 tasks involving generating HTML, LaTeX, SVG, and flowchart code from images, plus GUI reasoning. GPT-4o achieves the highest average among non-BigDocs models at 36.84, but all BigDocs-trained open models (49.52 to 56.34) outperform it. Claude-3.5 Sonnet scores 30.55, GeminiPro-1.5 scores 32.32, Qwen2-VL-72B scores 31.36, and Llama-3.2-90B scores 25.66. The gap is largest on flow generation and GUI reasoning tasks.

Evidence
correlational
Key metric
GPT-4o avg 36.84, Claude-3.5 Sonnet 30.55, GeminiPro-1.5 32.32, Qwen2-VL-72B 31.36, Llama-3.2-90B 25.66, Idefics2-8B 20.48; BigDocs-trained models 49.52-56.34
Caveat
BigDocs-Bench is a new benchmark introduced by this paper; the tasks may not be representative of general document understanding. Some tasks required 1-shot prompts for certain models (marked with dagger).
Model
GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 / Gemini Pro 1.5, Qwen2-VL Qwen2-VL-72B, Llama 3.2 Llama-3.2-90B, Idefics2 Idefics2-8B
Related findings
IC-344, IC-345
Extraction
automatic-extraction