IC-093No evaluated LLM achieves perfect confabulation avoidance on questions about non-existent events

Alexis Huet, Zied Ben Houidi, Dario Rossi

SourceEpisodic Memories Generation and Evaluation Benchmark for Large Language Models

The benchmark includes 150 questions whose cues match zero events in the document, testing whether the model can say 'I don't know.' No model reaches F1 of 1.0. O1-mini performs best at 0.97±0.16, followed by Claude 3.5 Sonnet at 0.92±0.27 (in-context) and GPT-4o at 0.84±0.37. Smaller models (GPT-4o-mini at 0.51±0.50) struggle substantially. Manual analysis of GPT-4o's 24 incorrect zero-event answers shows the model fabricates contextually relevant but factually wrong details by combining elements from different chapters.

Evidence
correlational
Key metric
O1-mini 0.97±0.16, Claude 3.5 Sonnet 0.92±0.27, GPT-4o 0.84±0.37, Claude 3 Haiku 0.84±0.37, Llama 3.1-405B 0.80±0.40, GPT-4o-mini 0.51±0.50 (in-context, long book, 150 zero-event questions)
Caveat
The zero-event questions use either inner (elements from the book but wrong combination) or outer (elements from outside the book) strategies; the authors note the evaluation uses a lenient F1 that may undercount false positives.
Model
GPT-4o mini, Claude 3 Haiku, Claude 3.5 Sonnet, Llama 3.1 405B Instruct, O1 / OpenAI-o1-preview O1-mini
Concepts
Failure mode
Datasets
Episodic Memory Benchmark / Episodic Memory Benchmark (short book) [eval]
Methods
In-Context Learning / In-context learning prompt [primary]
Related findings
IC-092, IC-094, IC-095
Extraction
automatic-extraction