Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
TabMWP
anchor
Findings
IC-1007
LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback
[eval]
IC-1008
The magnitude of CRITIC's improvement on mathematical program synthesis scales with Llama-2 model size
[eval]
IC-799
WizardMath-70b scores lower than base Llama-2-70b on TabMWP (49.8% vs 57.5%), indicating degraded OOD generalization from rationale-based fine-tuning
[eval]