Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
SVAMP
anchor
Findings
IC-1007
LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback
[eval]
IC-1008
The magnitude of CRITIC's improvement on mathematical program synthesis scales with Llama-2 model size
[eval]
IC-1264
LLMs are overconfident when verbalizing confidence, with values concentrated in 80–100% and multiples of 5, yielding high ECE across all five tested models
[eval]
IC-1265
Calibration and failure prediction improve as model capability scales from GPT-3 to GPT-4, but remain far from ideal
[eval]
IC-1266
For GPT-3, white-box token-probability methods outperform black-box verbalized confidence in uncertainty estimation, but the gap is narrow (0.522–0.605 AUROC) and both remain near random
[eval]