Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-312
GPT-4o achieves the highest insight-level Llama-3-eval score (0.60) among all tested LLM backbones on InsightBench multi-step data analytics
IC-313
All four LLM backbones fail to detect a planted linear trend in incident resolution time when the slope is below 0.1, and detection rates diverge sharply above that threshold
IC-314
Llama-3-70b as an LLM-based evaluator (Llama-3-eval) produces agent rankings consistent with GPT-4-based G-Eval on InsightBench