PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 20251 citationsOpen Access

VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation

View Full Paper
CZChaofan ZhangHPHao PengXCXin Cao

Key Points

  • The VTLA model outperforms traditional methods, achieving over 90% success rates on insertion tasks.
  • A low-cost, multi-modal dataset was constructed for training, focusing on vision-tactile-action-instruction pairs.
  • Direct Preference Optimization bridges continuous robotic tasks with classification-based models, enhancing performance.
  • Real-world experiments confirm the Sim2Real capability of the VTLA model, highlighting its practical application.

Abstract

While vision-language models have advanced significantly, their application in language-conditioned robotic manipulation is still underexplored, especially for contact-rich tasks that extend beyond visually dominant pick-and-place scenarios. To bridge this gap, we introduce Vision-Tactile-Language-Action model, a novel framework that enables robust policy generation in contact-intensive scenarios by effectively integrating visual and tactile inputs through cross-modal language grounding. A low-cost, multi-modal dataset has been constructed in a simulation environment, containing vision-tactile-action-instruction pairs specifically designed for the fingertip insertion task. Furthermore, we introduce Direct Preference Optimization (DPO) to offer regression-like supervision for the VTLA model, effectively bridging the gap between classification-based next token prediction loss and continuous robotic tasks. Experimental results show that the VTLA model outperforms traditional imitation learning methods (e.g., diffusion policies) and existing multi-modal baselines (TLA/VLA), achieving over 90% success rates on unseen peg shapes. Finally, we conduct real-world peg-in-hole experiments to demonstrate the exceptional Sim2Real performance of the proposed VTLA model. For supplementary videos and results, please visit our project website: https://sites.google.com/view/vtla

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68e8439a9989581a2fd4e07ahttps://doi.org/10.48550/arxiv.2505.09577
Ask AI
Helpful
Bookmark
Share
View Full Paper