Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-202
All six evaluated LLMs achieve very low accuracy on OpenRCA, with no model solving any three-element root cause query
IC-203
Gemini 1.5 Pro's RCA-Agent accuracy drops 68.4% when code execution fails, far exceeding the drops for Claude 3.5 (17.9%) and GPT-4o (15.6%)
IC-204
GPT-4o performs worse with explicit chain-of-thought prompting than with the original prompt on OpenRCA tasks