Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
VAL
anchor
Findings
IC-079
GPT-4's free-form critique generation is unreliable, containing hallucinated edges, vertex colors, and precondition states
[validation]
IC-080
GPT-4's performance is largely insensitive to the content of feedback; simple re-prompting with a sound verifier (sampling) matches or exceeds detailed critique
[validation]