Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-798
GPT-4 achieves 61.6% on MATH with tool-integrated reasoning prompting, exceeding PaL (51.8%) and CoT (42.5%)
IC-799
WizardMath-70b scores lower than base Llama-2-70b on TabMWP (49.8% vs 57.5%), indicating degraded OOD generalization from rationale-based fine-tuning