IC-074Released LLMs achieve F1 plan scores between 42.7 and 86.7 on the T-Eval plan task

Shengda Fan, Xin Cong, Yuepeng Fu, Zhong Zhang, Shuyan Zhang, Yuanwei Liu, Yesai Wu, Yankai Lin, Zhiyuan Liu, Maosong Sun

SourceWorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models

The paper reports F1 scores on the T-Eval plan task for a range of released models to contextualise WorkflowLlama's out-of-distribution generalisation. Proprietary models lead: GPT-4 scores 86.7, GPT-3.5 scores 86.6, Claude 2 scores 84.9. Open-source models range from Qwen-72B at 73.4 down to WizardLM-70B at 42.7. Llama-3.1-8B scores 68.2, Llama-2-70B scores 63.1, and Mistral-7B scores 64.9.

Evidence
correlational
Key metric
F1: GPT-4 86.7, GPT-3.5 86.6, Claude 2 84.9, Qwen-72B 73.4, Llama-3.1-8B 68.2, Llama-2-70B 63.1, Mistral-7B 64.9, WizardLM-70B 42.7
Caveat
Scores for baseline models may be taken from the original T-Eval paper rather than re-run by the authors; only WorkflowLlama was retrained and re-evaluated on the transformed T-Eval format.
Model
GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Qwen Qwen-7B, Qwen-14B, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3.1 8B, Llama 2 / Llama 2 base Llama 2 13B, Llama 2 70B, Vicuna Vicuna-13B, Baichuan2-13B, Wizardlm WizardLM-70B, Qwen1.5 Qwen-72B
Datasets
T-Eval [eval]
Methods
F1 score [eval]
Related work
T-Eval [context]
Related findings
IC-073, IC-075
Extraction
automatic-extraction