IC-1188GPT-3 procedural planning performance is scale-dependent, with Curie (6.7B) scoring 3.75 and Davinci (175B) scoring 4.90 overall quality in few-shot settings, while GPT-4 achieves 4.81 overall and 5.00 order

Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Lorraine Li, Hirona Jacqueline Arai, Soumya Sanyal, Keisuke Sakaguchi, Xiang Ren, Yejin Choi

SourcePlaSma: Procedural Knowledge Models for Language-based Planning and Re-Planning

The paper evaluates GPT-3 variants and GPT-4 on language-based procedural planning (decomposing goals into temporally ordered steps) using human evaluation on 250 goals. In a 5-shot setting, GPT-3 Curie achieves 3.75 overall quality while GPT-3 Davinci (175B) achieves 4.90, a 31% relative gap attributable to scale. GPT-4 in a few-shot setting on 50 instances achieves 4.81 overall and 5.00 order, comparable to Davinci. On constrained and counterfactual replanning (300 examples), Davinci zero-shot achieves 93.33% and 90% good plans respectively, while Curie zero-shot achieves only 68% and 44.33%.

Evidence
correlational
Key metric
GPT-3 Curie few-shot(5): 3.75 overall; GPT-3 Davinci zero-shot: 4.84, few-shot(5): 4.90 overall; GPT-4 few-shot: 4.81 overall, 5.00 order; CoCoGen few-shot(16): 4.55 overall; Constrained: Davinci zero-shot 93.33%, Curie zero-shot 68%; Counterfactual: Davinci zero-shot 90%, Curie zero-shot 44.33%
Model
GPT-3 / GPT base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Concepts
Scale-dependent behaviour
Datasets
ProScript [eval]
Extraction
automatic-extraction