Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
PlanBench
anchor
Findings
IC-078
GPT-4's self-verification loop causes performance collapse due to high false negative rates in binary verification
[eval]
IC-079
GPT-4's free-form critique generation is unreliable, containing hallucinated edges, vertex colors, and precondition states
[eval]
IC-080
GPT-4's performance is largely insensitive to the content of feedback; simple re-prompting with a sound verifier (sampling) matches or exceeds detailed critique
[eval]
IC-549
All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBench
[compared-to]