PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 21, 20240 citationsOpen Access

Open-vocabulary Pick and Place via Patch-level Semantic Maps

View Full Paper
MJMingxi JiaXinjiang Agricultural UniversityHHHaojie HuangPioneer (United States)ZZZhewen ZhangFudan University

Key Points

Key points are not available for this paper at this time.

Abstract

Controlling robots through natural language instructions in open-vocabulary scenarios is pivotal for enhancing human-robot collaboration and complex robot behavior synthesis. However, achieving this capability poses significant challenges due to the need for a system that can generalize from limited data to a wide range of tasks and environments. Existing methods rely on large, costly datasets and struggle with generalization. This paper introduces Grounded Equivariant Manipulation (GEM), a novel approach that leverages the generative capabilities of pre-trained vision-language models and geometric symmetries to facilitate few-shot and zero-shot learning for open-vocabulary robot manipulation tasks. Our experiments demonstrate GEM's high sample efficiency and superior generalization across diverse pick-and-place tasks in both simulation and real-world experiments, showcasing its ability to adapt to novel instructions and unseen objects with minimal data requirements. GEM advances a significant step forward in the domain of language-conditioned robot control, bridging the gap between semantic understanding and action generation in robotic systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jia et al. (2024) studied this question.

synapsesocial.com/papers/68e63e20b6db6435875cfc1ahttps://doi.org/10.48550/arxiv.2406.15677
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-Guided 3D Policy2025 · 11 citations
  2. 2Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps2024
  3. 3MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting2024 · 5 citations
  4. 4Manipulate-Anything: Automating Real-World Robots using Vision-Language Models2024 · 5 citations
  5. 5PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation2025