The paper reports F1 scores on the T-Eval plan task for a range of released models to contextualise WorkflowLlama's out-of-distribution generalisation. Proprietary models lead: GPT-4 scores 86.7, GPT-3.5 scores 86.6, Claude 2 scores 84.9. Open-source models range from Qwen-72B at 73.4 down to WizardLM-70B at 42.7. Llama-3.1-8B scores 68.2, Llama-2-70B scores 63.1, and Mistral-7B scores 64.9.
Scores for baseline models may be taken from the original T-Eval paper rather than re-run by the authors; only WorkflowLlama was retrained and re-evaluated on the transformed T-Eval format.