PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 20240 citationsOpen Access

Driving Referring Video Object Segmentation with Vision-Language Pre-trained Models

View Full Paper
ZZZikun ZhouWXWentao XiongLZLi Zhou

Key Points

  • The VLP-RVOS framework outperforms existing state-of-the-art algorithms across referring video object segmentation benchmarks, demonstrating superior generalization abilities.
  • Temporal-aware prompt-tuning adapts pre-trained vision-language representations to dynamic video clues, enabling accurate pixel-level prediction alongside cube-frame attention.
  • Multi-stage relation modeling unifies cross-modal feature extraction and spatial-temporal reasoning, providing an effective paradigm to transfer vision-language pre-trained models.

Abstract

The crux of Referring Video Object Segmentation (RVOS) lies in modeling dense text-video relations to associate abstract linguistic concepts with dynamic visual contents at pixel-level. Current RVOS methods typically use vision and language models pre-trained independently as backbones. As images and texts are mapped to uncoupled feature spaces, they face the arduous task of learning Vision-Language~(VL) relation modeling from scratch. Witnessing the success of Vision-Language Pre-trained (VLP) models, we propose to learn relation modeling for RVOS based on their aligned VL feature space. Nevertheless, transferring VLP models to RVOS is a deceptively challenging task due to the substantial gap between the pre-training task (image/region-level prediction) and the RVOS task (pixel-level prediction in videos). In this work, we introduce a framework named VLP-RVOS to address this transfer challenge. We first propose a temporal-aware prompt-tuning method, which not only adapts pre-trained representations for pixel-level prediction but also empowers the vision encoder to model temporal clues. We further propose to perform multi-stage VL relation modeling while and after feature extraction for comprehensive VL understanding. Besides, we customize a cube-frame attention mechanism for spatial-temporal reasoning. Extensive experiments demonstrate that our method outperforms state-of-the-art algorithms and exhibits strong generalization abilities.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhou et al. (2024) studied this question.

synapsesocial.com/papers/68e69ae8b6db6435876205abhttps://doi.org/10.48550/arxiv.2405.10610
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation2025
  2. 2Cross-modal Object Decoding and Referring Expression Decoupling for Referring Video Object Segmentation2024
  3. 3UNINEXT-Cutie: The 1st Solution for LSVOS Challenge RVOS Track2024
  4. 4SVAC: Scaling Is All You Need For Referring Video Object Segmentation2025
  5. 5Language as Queries for Referring Video Object Segmentation2022 · 166 citations