IC-551GPT-4's workflow generation performance declines as the number of nodes and edges in the workflow increases

Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen

SourceBenchmarking Agentic Workflow Generation

The paper analyzes GPT-4's F1_chain and F1_graph scores binned by the number of nodes and edges in the gold workflow (Figure 3). As workflow complexity increases from 2 to 12+ nodes/edges, both metrics show a general downward trend with occasional brief spikes attributed to uneven sample distribution. The authors conclude that for complex planning tasks with more steps, GPT-4's performance is unsatisfying for both linear and graph planning.

Evidence
correlational
Caveat
The specific per-bin values are shown only in Figure 3 and are not printed as text or table values. The trend is described qualitatively in the text. Occasional spikes are attributed to uneven sample distribution.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Concepts
Failure mode
Related findings
IC-549, IC-550, IC-552
Extraction
automatic-extraction