PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 20250 citationsOpen Access

CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision

View Full Paper
YLYifei LuFYFanghua YeJLJian Li

Key Points

  • CodeTool significantly improves tool invocation in large language models, providing immediate feedback on each step.
  • Extensive experiments on StableToolBench and RestBench-TMDB highlight the effectiveness of this new approach.
  • This framework leverages distinct process rewards, enhancing the overall task completion process for LLMs.
  • Maximizing the cumulative rewards guides LLMs toward efficient and accurate reasoning paths.

Abstract

Tool invocation significantly enhances the capabilities of Large Language Models (LLMs), yet challenges persist, particularly in complex task scenarios. Current methods, such as instruction-enhanced reasoning and supervised fine-tuning, often result in unnecessarily long reasoning paths and face difficulties in verifying the correctness of intermediate steps. In this paper, we propose CodeTool, a novel framework for stepwise code generation that improves LLM tool invocation by leveraging the concise and easily verifiable nature of code. CodeTool incorporates two distinct process rewards: the On-the-spot Reward, which provides immediate feedback on the accuracy of each tool invocation, and the Latent Reward, which assesses the contribution of each step toward overall task completion. By maximizing the cumulative reward of the On-the-spot and Latend Rewards at each step, LLMs are guided to follow efficient and accurate reasoning paths. Extensive experiments on StableToolBench and RestBench-TMDB demonstrate the superiority of CodeTool over existing approaches.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lu et al. (2025) studied this question.

synapsesocial.com/papers/68ec384042a190b2c3519a5chttps://doi.org/10.48550/arxiv.2503.20840
Ask AI
Helpful
Bookmark
Share
View Full Paper