PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 20250 citationsOpen Access

Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation

View Full Paper
SLShunlei LiLGLongsen GaoJCJiuwen Cao

Key Points

  • The framework enables dual-arm robots to achieve over 95% accuracy in task execution and reasoning.
  • Using temporal scene graphs and language conditioning, GF-VLA enhances robotic performance in complex tasks.
  • Cross-arm allocation strategy autonomously assigns grippers, improving bimanual execution efficiency.
  • Robust policies generated by GF-VLA demonstrated high grasp reliability and task success rates in various scenarios.

Abstract

Acquiring dexterous robotic skills from human video demonstrations remains a significant challenge, largely due to conventional reliance on low-level trajectory replication, which often fails to generalize across varying objects, spatial layouts, and manipulator configurations. To address this limitation, we introduce Graph-Fused Vision-Language-Action (GF-VLA), a unified framework that enables dual-arm robotic systems to perform task-level reasoning and execution directly from RGB-D human demonstrations. GF-VLA employs an information-theoretic approach to extract task-relevant cues, selectively highlighting critical hand-object and object-object interactions. These cues are structured into temporally ordered scene graphs, which are subsequently integrated with a language-conditioned transformer to produce hierarchical behavior trees and interpretable Cartesian motion primitives. To enhance efficiency in bimanual execution, we propose a cross-arm allocation strategy that autonomously determines gripper assignment without requiring explicit geometric modeling. We validate GF-VLA on four dual-arm block assembly benchmarks involving symbolic structure construction and spatial generalization. Empirical results demonstrate that the proposed representation achieves over 95% graph accuracy and 93% subtask segmentation, enabling the language-action planner to generate robust, interpretable task policies. When deployed on a dual-arm robot, these policies attain 94% grasp reliability, 89% placement accuracy, and 90% overall task success across stacking, letter-formation, and geometric reconfiguration tasks, evidencing strong generalization and robustness under diverse spatial and semantic variations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68ec384042a190b2c351999bhttps://doi.org/10.48550/arxiv.2509.07957
Ask AI
Helpful
Bookmark
Share
View Full Paper