IC-312GPT-4o achieves the highest insight-level Llama-3-eval score (0.60) among all tested LLM backbones on InsightBench multi-step data analytics

Gaurav Sahu, Abhay Puri, Juan A. Rodriguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, Nicolas Chapados, Christopher Pal, Sai Rajeswar, Issam H. Laradji

SourceInsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation

The paper evaluates four released LLMs as backbones in the AgentPoirot data analytics agent on 100 synthetic business datasets. GPT-4o attains the best insight-level Llama-3-eval score of 0.60 (±0.03), followed by GPT-4-turbo (0.56±0.02), Llama-3-70b (0.52±0.04), and GPT-3.5-turbo (0.50±0.02). The Pandas Agent baseline with GPT-4o scores 0.54±0.01 on the same metric. Performance also varies by dataset category, difficulty, and insight type, with all agents showing declining performance on harder tasks and from descriptive to predictive insights.

Evidence
correlational
Key metric
insight-level llama-3-eval: GPT-4o 0.60±0.03, GPT-4-turbo 0.56±0.02, Llama-3-70b 0.52±0.04, GPT-3.5-turbo 0.50±0.02 (AgentPoirot); PA (GPT-4o) 0.54±0.01
Caveat
Results are measured through a specific agent framework (AgentPoirot) with fixed prompting; performance may differ under other agent designs. All results are for sampling temperature 0.0 and 5 seeds.
Model
GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Llama 3 70B
Datasets
InsightBench [eval]
Methods
ROUGE-1 [eval]
Related work
G-Eval [context]
Related findings
IC-313, IC-314
Extraction
automatic-extraction