PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

View Full Paper
SKSunil KumarBZBowen ZhaoLDLeo Dirac

Key Points

  • Our method improves visual reasoning capabilities of vision-language models, especially under resource constraints.
  • Using Group Relative Policy Optimization with a simple reward structure enhances performance on visual question-answering tasks.
  • Models trained on a diverse data mix show superior outcomes, leveraging external tools for detailed visual information.
  • This approach highlights potential for advancing smaller models without significant computational overhead.

Abstract

Despite tremendous recent advances in large model reasoning ability, vision-language models (VLMs) still struggle with detailed visual reasoning, especially when compute resources are limited. To address this challenge, we draw inspiration from methods like Deepseek-r1 for VLMs and train smaller-scale models with Group Relative Policy Optimization (GRPO) to use external tools such as zoom. The greatest benefit is obtained with a combination of GRPO learning, a simple reward structure, a simplified tool-calling interface, allocating additional tokens to the result of the tool call, and a training data mix that over-represents visually difficult examples. Compared to similarly-sized baseline models, our method achieves better performance on some visual question-answering (VQA) tasks, thanks to the detailed visual information gathered from the external tool.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kumar et al. (2025) studied this question.

synapsesocial.com/papers/68f6379bb481a140a36cf54bhttps://doi.org/10.48550/arxiv.2506.14821
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning2026
  2. 2Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning2025 · 1 citations
  3. 3VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use2025
  4. 4Enhancing Large Vision-Language Models via Quantized Grounded Reasoning2025
  5. 5Empowering Vision-Language Models for Reasoning Ability through Large Language Models2024 · 8 citations