PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 3, 20250 citationsOpen Access

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

View Full Paper
YZYang ZhangCWChenwei WangŎO ̈sürmeliog ̆lu

Key Points

  • Align-Then-stEer enhances vision-language-action models by improving adaptation for robotic tasks.
  • The framework leads to a 9.8% success rate increase in simulation and 32% in real-world settings.
  • By creating a unified latent space, it effectively aligns different action distributions for fine-tuning.
  • The use of a variational autoencoder and guidance mechanism streamlines the adaptation process for diverse robotics applications.

Abstract

Vision-Language-Action (VLA) models pre-trained on large, diverse datasets show remarkable potential for general-purpose robotic manipulation. However, a primary bottleneck remains in adapting these models to downstream tasks, especially when the robot's embodiment or the task itself differs from the pre-training data. This discrepancy leads to a significant mismatch in action distributions, demanding extensive data and compute for effective fine-tuning. To address this challenge, we introduce Align-Then-stEer (ATE), a novel, data-efficient, and plug-and-play adaptation framework. ATE first aligns disparate action spaces by constructing a unified latent space, where a variational autoencoder constrained by reverse KL divergence embeds adaptation actions into modes of the pre-training action latent distribution. Subsequently, it steers the diffusion- or flow-based VLA's generation process during fine-tuning via a guidance mechanism that pushes the model's output distribution towards the target domain. We conduct extensive experiments on cross-embodiment and cross-task manipulation in both simulation and real world. Compared to direct fine-tuning of representative VLAs, our method improves the average multi-task success rate by up to 9. 8\% in simulation and achieves a striking 32\% success rate gain in a real-world cross-embodiment setting. Our work presents a general and lightweight solution that greatly enhances the practicality of deploying VLA models to new robotic platforms and tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68e02f40f0e39f13e7fa28dahttps://doi.org/10.48550/arxiv.2509.02055
Ask AI
Helpful
Bookmark
Share
View Full Paper