Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MuSiQue / Musique-ans
anchor
Findings
IC-1598
Retrieval augmentation improves GPT-3.5-turbo-4k on long-context tasks but not GPT-3.5-turbo-16k
[eval]
IC-174
RAG reduces model abstention and LLMs hallucinate rather than abstain when the retrieved context is insufficient to answer the query
[eval]
IC-175
Context-sufficiency performance is scale-dependent: larger LLMs achieve high accuracy with sufficient context but still answer correctly 35-62% of the time without it, while smaller models hallucinate or abstain even with sufficient context
[eval]
IC-485
LLMs show a significant performance gap between Wikipedia-based factual multi-hop QA and counterfactual multi-hop QA, indicating reliance on memorized knowledge rather than reasoning from context
[eval]