Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CRITIC
anchor
Findings
IC-1007
LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback
[primary]
IC-1008
The magnitude of CRITIC's improvement on mathematical program synthesis scales with Llama-2 model size
[primary]