PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks

View Full Paper
EHEric HannusMMMartha MalinTLTran Nguyen Le

Key Points

  • IA-VLA enhances the action-output capabilities of vision-language-action models, especially in semantically complex tasks.
  • Experiments reveal that augmenting input with context significantly aids the model’s performance when addressing language instructions.
  • The evaluation used a dataset of scenes with visually indistinguishable objects to rigorously compare VLA configurations.
  • Findings highlight the potential of input augmentation in overcoming challenges in robot manipulation and language comprehension.

Abstract

Vision-language-action models (VLAs) have become an increasingly popular approach for addressing robot manipulation problems in recent years. However, such models need to output actions at a rate suitable for robot control, which limits the size of the language model they can be based on, and consequently, their language understanding capabilities. Manipulation tasks may require complex language instructions, such as identifying target objects by their relative positions, to specify human intention. Therefore, we introduce IA-VLA, a framework that utilizes the extensive language understanding of a large vision language model as a pre-processing stage to generate improved context to augment the input of a VLA. We evaluate the framework on a set of semantically complex tasks which have been underexplored in VLA literature, namely tasks involving visual duplicates, i.e., visually indistinguishable objects. A dataset of three types of scenes with duplicate objects is used to compare a baseline VLA against two augmented variants. The experiments show that the VLA benefits from the augmentation scheme, especially when faced with language instructions that require the VLA to extrapolate from concepts it has seen in the demonstrations. For the code, dataset, and videos, see https://sites.google.com/view/ia-vla.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hannus et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcce8d54a28a75cf1d3dhttps://doi.org/10.48550/arxiv.2509.24768
Ask AI
Helpful
Bookmark
Share
View Full Paper