Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Teaching Large Language Models to Self-Debug
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-897
GPT-4 underperforms codex on Spider text-to-SQL with few-shot prompting, attributed to its zero-shot tuning
IC-898
GPT-3.5-turbo and GPT-4 are overconfident in their initial code predictions when unit test execution is unavailable