Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Benchmarking Agentic Workflow Generation
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-549
All 18 evaluated LLMs show a 15-20% performance gap between linear (node chain) and graph (workflow) planning on WorfBench
IC-550
Workflow generation performance scales with model size within families, but recently released 7B models outperform older 13B models
IC-551
GPT-4's workflow generation performance declines as the number of nodes and edges in the workflow increases
IC-552
GPT-4, Llama-3.1-8B, and Qwen-2-72B all improve on ALFWorld and WebShop when given a generated workflow as structured prior knowledge