Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-485
LLMs show a significant performance gap between Wikipedia-based factual multi-hop QA and counterfactual multi-hop QA, indicating reliance on memorized knowledge rather than reasoning from context
IC-486
LLMs achieve correct final answers through incorrect reasoning chains, inflating their apparent multi-step reasoning performance
IC-487
Including sub-questions in the prompt improves LLM performance on multi-hop QA tasks
IC-488
LLM performance degrades progressively as the number of reasoning hops increases, with error propagation from earlier sub-questions