IC-989Tool selection CSR degrades as the candidate tool list grows from 5 to 15 tools, and performance varies by user scenario

Yue Huang, Jiawen Shi, Yuan Li, Chenrui Fan, Siyuan Wu, Qihui Zhang, Yixin Liu, Pan Zhou, Yao Wan, Neil Zhenqiang Gong, Lichao Sun

SourceMetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Across all eight LLMs, the correct selection rate drops as the number of candidate tools increases from top-5 to top-10 to top-15, with the steepest decline between top-5 and top-10. For example, ChatGPT zero-shot CSR falls from 88.89 (top-5) to 77.39 (top-10) to 75.08 (top-15); Vicuna-33b drops from 92.00 to 73.23 to 70.90. Performance also varies by user scenario: LLMs generally score highest for elderly and artists-and-designers tool sets and lowest for student-related tools, indicating a domain-dependent bias in tool selection.

Evidence
correlational
Key metric
Zero-shot CSR by list size (popularity scenarios): ChatGPT 88.89/77.39/75.08 (top5/10/15); Vicuna-33b 92.00/73.23/70.90; ChatGLM2 82.29/56.19/48.07; Llama2-13b 42.00/43.72/38.93.
Caveat
The scenario-based tool lists are manually curated (10 tools per scenario), and the popularity-based lists are ranked by the number of merged tools, which may not reflect real-world tool popularity.
Model
ChatGPT, ChatGLM2, Llama 2 / Llama 2 base, Vicuna, Koala
Concepts
Failure mode
Related work
API-Bank [compared-to], ToolLLM [compared-to]
Related findings
IC-986, IC-987, IC-988
Extraction
automatic-extraction