Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1144
GPT-4 and GPT-3.5 produce high rates of irrelevant (fabricated) books when answering constraint queries from parametric knowledge, with a sharp phase transition at low author popularity
IC-1145
Providing complete context eliminates irrelevance but does not fix constraint satisfaction for GPT-4 or GPT-3.5
IC-1146
Self-context (self-retrieval) chain-of-thought increases the rate of fabricated books compared to no-context for both GPT-4 and GPT-3.5
IC-1147
GPT-4 outperforms GPT-3.5 on all KITAB metrics but the gap is modest, with all-correctness below 35% for both, suggesting scale alone does not resolve constraint satisfaction