Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
HaDeMiF: Hallucination Detection and Mitigation in Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-282
GPT-2 XL (1.5B) exhibits lower accuracy but reduced overconfidence (smaller ECE and Brier scores) compared to larger models on the CAT benchmark