PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 28, 20240 citationsOpen Access

TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

View Full Paper
BJBu JinYZYupeng ZhengPLPengfei Li

Key Points

Key points are not available for this paper at this time.

Abstract

3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D dense captioning in outdoor scenes is hindered by two major challenges: 1) the domain gap between indoor and outdoor scenes, such as dynamics and sparse visual inputs, makes it difficult to directly adapt existing indoor methods; 2) the lack of data with comprehensive box-caption pair annotations specifically tailored for outdoor scenes. To this end, we introduce the new task of outdoor 3D dense captioning. As input, we assume a LiDAR point cloud and a set of RGB images captured by the panoramic camera rig. The expected output is a set of object boxes with captions. To tackle this task, we propose the TOD3Cap network, which leverages the BEV representation to generate object box proposals and integrates Relation Q-Former with LLaMA-Adapter to generate rich captions for these objects. We also introduce the TOD3Cap dataset, the largest one to our knowledge for 3D dense captioning in outdoor scenes, which contains 2. 3M descriptions of 64. 3K outdoor objects from 850 scenes. Notably, our TOD3Cap network can effectively localize and caption 3D objects in outdoor scenes, which outperforms baseline methods by a significant margin (+9. 6 CiDEr@0. 5IoU). Code, data, and models are publicly available at https: //github. com/jxbbb/TOD3Cap.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jin et al. (2024) studied this question.

synapsesocial.com/papers/68e720d3b6db64358769a66chttps://doi.org/10.48550/arxiv.2403.19589
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioning2024 · 24 citations
  2. 2Lowis3D: Language-Driven Open-World Instance-Level 3D Scene Understanding2024 · 26 citations
  3. 3View Selection for 3D Captioning via Diffusion Ranking2024
  4. 4Bi-directional Contextual Attention for 3D Dense Captioning2024
  5. 5Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization2024