Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1327
GPT-4, GPT-3.5, Llama2, and Vicuna models underperform human annotators on multistep soft reasoning in natural language narratives, with smaller models scoring near random chance
IC-1328
GPT-4's performance on MUSR depends on the type of neurosymbolic scaffolding: program-aided decomposition helps on structured optimization but symbolic belief tracking fails on natural language theory-of-mind