Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Game of 24 (4nums.com instances 1-1000)
Findings
IC-078
GPT-4's self-verification loop causes performance collapse due to high false negative rates in binary verification
[eval]
IC-079
GPT-4's free-form critique generation is unreliable, containing hallucinated edges, vertex colors, and precondition states
[eval]
IC-080
GPT-4's performance is largely insensitive to the content of feedback; simple re-prompting with a sound verifier (sampling) matches or exceeds detailed critique
[eval]