Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-092
All evaluated LLMs show consistent F1 degradation to at most 0.60 when two or more events match a retrieval cue
IC-093
No evaluated LLM achieves perfect confabulation avoidance on questions about non-existent events
IC-094
Episodic recall accuracy degrades systematically from content cues to space cues to time cues across all evaluated LLMs
IC-095
Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological ordering