PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 10, 20260 citationsOpen Access

CompToolBench: A Compositional Benchmark for Fine-Grained Tool-Use Evaluation of Large Language Models

View Full Paper
MRMd A Rahman

Key Points

  • This research aims to evaluate the performance of large language models (LLMs) in using tools across various compositional complexities.
  • Introduced CompToolBench as a benchmark for tool-use evaluation.
  • Assessed 18 LLMs drawn from cloud and local deployments.
  • Evaluated 200 tasks using 106 public APIs across four composition levels.
  • Identified performance gaps in tool selection and call validity.
  • Found an average Selection Gap of 13.2 percentage points in tool call validity.
  • Compositional complexity does not consistently degrade model performance.
  • Local models are now nearing the accuracy of cloud models on structured tasks.

Abstract

We introduce CompToolBench, a benchmark for evaluating LLM tool-use across four composition levels (single, sequential, parallel, and graph) using 200 tasks and 106 real tools drawn from free public APIs. We evaluate 18 models spanning cloud and local deployments and identify a Selection Gap: models that correctly select tools frequently fail to produce valid calls, with an average gap of 13.2 percentage points. Results show that compositional complexity does not uniformly degrade performance, and that local models now approach cloud-model accuracy on structured tool-use tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Md A Rahman (2026) studied this question.

synapsesocial.com/papers/69af95b470916d39fea4d7e5https://doi.org/10.5281/zenodo.18907783
Ask AI
Helpful
Bookmark
Share
View Full Paper