PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 20250 citationsOpen Access

VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback

View Full Paper
JBJianxin BiKMKevin MaCHCe Hao

Key Points

  • Integration of tactile feedback improves task planning efficiency and execution precision in robotics.
  • VLA-Touch relies on a pretrained tactile-language model to enhance robot task planning.
  • Utilizing a diffusion-based controller, VLA-Touch refines actions generated by VLA with tactile signals.
  • The method addresses the lack of large multi-modal datasets for incorporating tactile signals into VLA models.

Abstract

Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the ability to interpret and use tactile signals, limiting their effectiveness in contact-rich tasks. Incorporating tactile feedback into these systems is challenging due to the absence of large multi-modal datasets. We present VLA-Touch, an approach that enhances generalist robot policies with tactile sensing without fine-tuning the base VLA. Our method introduces two key innovations: (1) a pipeline that leverages a pretrained tactile-language model that provides semantic tactile feedback for high-level task planning, and (2) a diffusion-based controller that refines VLA-generated actions with tactile signals for contact-rich manipulation. Through real-world experiments, we demonstrate that our dual-level integration of tactile feedback improves task planning efficiency while enhancing execution precision. Code is open-sourced at https: //github. com/jxbi1010/VLA-Touchthis URL.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bi et al. (2025) studied this question.

synapsesocial.com/papers/68e6679587ecc93a24d1757ehttps://doi.org/10.48550/arxiv.2507.17294
Ask AI
Helpful
Bookmark
Share
View Full Paper