Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-610
In LLaMA2-7B-Chat, RAG hallucinations are causally driven by copying heads losing external context information during generation and by knowledge FFNs in mid-to-upper layers over-adding parametric knowledge to the residual stream