IC-345GPT-4o's table2latex outputs lose 63% of the time in human evaluation, with inconsistent formatting (lines, borders, margins) as the primary failure

Juan A. Rodriguez, Xiangru Jian, Siba Smarak Panigrahi, Tianyu Zhang, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Zichao Li, Suyuchen Wang, Pierre-Andre Noel, Mats Leon Richter, Saverio Vadacchino, Shubham Agarwal, Sanket Biswas, Sara Shanian, Ying Zhang, Sathwik Tejaswi Madhusudhan, Joao Monteiro, Krishnamurthy Dj Dvijotham, Torsten Scholak, Nicolas Chapados, Sepideh Kharaghani, Sean Hughes, M. Özsu, Siva Reddy, Marco Pedersoli, Yoshua Bengio, Christopher Pal, Issam H. Laradji, Spandana Gella, Perouz Taslakian, David Vazquez, Sai Rajeswar

SourceBigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks

In a human evaluation with 28 evaluators providing 1,900 annotations, GPT-4o's LaTeX table outputs were compared against phi3.5-vision+bigdocs. GPT-4o loses 63% of the time with 31% draws and 6% both bad. The authors note that GPT-4o 'often struggles to maintain consistent formatting despite capturing content accurately.' On screenshot2html, GPT-4o performs more competitively, with phi3.5+bigdocs winning only 36% and GPT-4o winning 48%.

Evidence
correlational
Key metric
table2latex: phi3.5 bigdocs win 63%, draw 31%, both bad 6%; screenshot2html: phi3.5 bigdocs win 36%, loss 48%, draw 7%, both bad 9%
Caveat
The comparison is against the authors' own BigDocs-trained model, which may be optimised for this specific task. The 8192-token context limit may affect GPT-4o's ability to generate complete HTML.
Model
GPT-4o
Concepts
Failure mode
Related findings
IC-343, IC-344
Extraction
automatic-extraction