IC-312GPT-4o achieves the highest insight-level Llama-3-eval score (0.60) among all tested LLM backbones on InsightBench multi-step data analytics
Gaurav Sahu, Abhay Puri, Juan A. Rodriguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, Nicolas Chapados, Christopher Pal, Sai Rajeswar, Issam H. Laradji
The paper evaluates four released LLMs as backbones in the AgentPoirot data analytics agent on 100 synthetic business datasets. GPT-4o attains the best insight-level Llama-3-eval score of 0.60 (±0.03), followed by GPT-4-turbo (0.56±0.02), Llama-3-70b (0.52±0.04), and GPT-3.5-turbo (0.50±0.02). The Pandas Agent baseline with GPT-4o scores 0.54±0.01 on the same metric. Performance also varies by dataset category, difficulty, and insight type, with all agents showing declining performance on harder tasks and from descriptive to predictive insights.
Results are measured through a specific agent framework (AgentPoirot) with fixed prompting; performance may differ under other agent designs. All results are for sampling temperature 0.0 and 5 seeds.