Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
APPS
anchor
Findings
IC-1231
On APPS, released code generation models span pass@1 from 0.20 (GPT-3 175B) to 6.20 (CodeRL), with value-based and policy-based RL methods outperforming supervised baselines
[eval]
IC-1599
Self-repair at equivalent compute budget provides only modest and inconsistent gains over i.i.d. sampling for CodeLlama-13B-Instruct, GPT-3.5, and GPT-4 on HumanEval and APPS
[eval]
IC-1600
Replacing a model's self-generated feedback with a stronger model's feedback consistently improves self-repair beyond both the i.i.d. baseline and the self-repair baseline
[eval]
IC-1601
GPT-4's self-generated feedback is significantly less effective than human programmer feedback for code repair, with the gap widening on harder problems
[eval]