IC-313All four LLM backbones fail to detect a planted linear trend in incident resolution time when the slope is below 0.1, and detection rates diverge sharply above that threshold

Gaurav Sahu, Abhay Puri, Juan A. Rodriguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, Nicolas Chapados, Christopher Pal, Sai Rajeswar, Issam H. Laradji

SourceInsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation

The paper varies the slope of a planted linear trend in dataset 2 from -0.1 to 0.9 and measures whether each LLM backbone, within AgentPoirot, discovers the insight. All models (GPT-4o, GPT-4-turbo, GPT-3.5-turbo, Llama-3-70b) fail to detect the trend when the slope is below 0.1, including negative slopes, with no false positives. Above slope 0.1, GPT-4o and GPT-4-turbo discover the insight every time, Llama-3-70b discovers it nearly half the time, and GPT-3.5-turbo has the lowest and most variable detection rate.

Evidence
correlational
Key metric
All models fail at slope < 0.1; for slope > 0.1: GPT-4o and GPT-4-turbo detect every time, Llama-3-70b nearly half the time, GPT-3.5-turbo particularly low detection rate
Caveat
Tested on a single dataset (dataset 2) with a single planted insight type (linear trend in TTR); generalizability to other trend types or datasets is not established.
Model
GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Llama 3 70B
Concepts
Failure mode
Datasets
InsightBench [eval]
Related findings
IC-312, IC-314
Extraction
automatic-extraction