IC-1188GPT-3 procedural planning performance is scale-dependent, with Curie (6.7B) scoring 3.75 and Davinci (175B) scoring 4.90 overall quality in few-shot settings, while GPT-4 achieves 4.81 overall and 5.00 order
The paper evaluates GPT-3 variants and GPT-4 on language-based procedural planning (decomposing goals into temporally ordered steps) using human evaluation on 250 goals. In a 5-shot setting, GPT-3 Curie achieves 3.75 overall quality while GPT-3 Davinci (175B) achieves 4.90, a 31% relative gap attributable to scale. GPT-4 in a few-shot setting on 50 instances achieves 4.81 overall and 5.00 order, comparable to Davinci. On constrained and counterfactual replanning (300 examples), Davinci zero-shot achieves 93.33% and 90% good plans respectively, while Curie zero-shot achieves only 68% and 44.33%.