PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20251 citationsOpen Access

A Neurosymbolic Agent System for Compositional Visual Reasoning

View Full Paper
YXYang XuGLGaowen LiuRKRamana Rao Kompella

Key Points

  • VLAgent achieves superior performance in compositional visual reasoning tasks compared to existing models.
  • Extensive experiments on visual benchmarks demonstrate that VLAgent outperforms over a dozen state-of-the-art visual reasoning systems.
  • The system integrates a front-end engine for generating structured reasoning plans and a back-end engine for executing these plans effectively.
  • VLAgent's unique features enhance quality through syntax checking, error detection, and validation processes during reasoning.

Abstract

The advancement in large language models (LLMs) and large vision models has fueled the rapid progress in multi-modal vision-language reasoning capabilities. However, existing vision-language models (VLMs) remain challenged by compositional visual reasoning. This paper presents VLAgent, a neuro-symbolic approach to developing a Vision-Language Agent system for efficient compositional visual reasoning with three novel features. First, VLAgent develops an interpretable visualization-enhanced two-stage neuro-symbolic reasoning system. The first stage is managed by a front-end engine that generates a structured visual reasoning plan (symbolic program script) for each compositional visual reasoning task by utilizing a pre-trained LLM powered with few-shot chain-of-thought in-context learning. The second stage is managed by a high-performance back-end engine. It transforms the planning script into executable code based on visual input (image or video) and the combination of neural models and symbolic functions and then performs a sequence of actions for the compositional visual reason task. Second, to ensure and enhance the quality of mapping the logic plan to a sequence of executable instructions, VLAgent introduces the SS-parser, which examines the syntax and semantic correctness of the planning script, detects and repairs the logic errors found in the LLM-generated logic plan before generating the executable program. Third, VLAgent introduces the execution verifier in critical reasoning steps to validate and refine its compositional reasoning results in a stepwise manner, for example, ensemble methods for critical visual reasoning and caption analysis for low-confidence compositional reasoning. Extensive experiments on six visual benchmarks compared to a dozen SoTA visual reasoning models show that VLAgent outperforms existing representative approaches to compositional visual reasoning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xu et al. (2025) studied this question.

synapsesocial.com/papers/68e861a57ef2f04ca37e47achttps://doi.org/10.48550/arxiv.2506.07778
Ask AI
Helpful
Bookmark
Share
View Full Paper