Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-343
GPT-4o, Claude-3.5 Sonnet, and GeminiPro-1.5 score below BigDocs-trained open models on BigDocs-Bench tasks requiring long structured code generation
IC-344
GPT-4o achieves the highest average score (64.62) on general document benchmarks, outperforming Qwen2-VL-72B (58.40) and GeminiPro-1.5 (57.05)
IC-345
GPT-4o's table2latex outputs lose 63% of the time in human evaluation, with inconsistent formatting (lines, borders, margins) as the primary failure