Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-986
Most LLMs lack tool usage awareness, with only ChatGPT exceeding 70% F1 in zero-shot evaluation
IC-987
When the correct tool is absent from the candidate list, most LLMs hallucinate a tool rather than returning 'none'
IC-988
LLMs show large gaps in multi-tool selection and over-rely on the number of tools specified in the prompt
IC-989
Tool selection CSR degrades as the candidate tool list grows from 5 to 15 tools, and performance varies by user scenario