Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Retrieval-Augmented Generation / Retrieval augmentation (top-5 chunks)
anchor
Findings
IC-092
All evaluated LLMs show consistent F1 degradation to at most 0.60 when two or more events match a retrieval cue
[primary]
IC-094
Episodic recall accuracy degrades systematically from content cues to space cues to time cues across all evaluated LLMs
[primary]
IC-095
Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological ordering
[primary]
IC-1035
Reprover achieves 0% accuracy on sorry theorems in advanced mathematics repositories (PFR, Hairy Ball Theorem, Coxeter) while proving basic theorems in other repositories
[supporting]
IC-1598
Retrieval augmentation improves GPT-3.5-turbo-4k on long-context tasks but not GPT-3.5-turbo-16k
[primary]
IC-555
Large LLMs (GPT-3.5-turbo, Gemini 1.5 Flash, Llama3-70B, Mixtral 46.7B) exhibit reasoning errors and significant accuracy degradation on large-scale logical commonsense reasoning tasks with 32k+ rules, even when the knowledge base is complete and retrieval is ideal
[primary]