PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 8, 20240 citations

Open-Vocabulary And Multitask Image Segmentation

View Full Paper
LPLihu PanYYYunting YangZWZhengkui Wang

Key Points

  • OVAMTSeg achieves a 51.6 mean Intersection over Union (mIoU) on Pascal-VOC with four unseen classes, showing marked effectiveness.
  • Key metrics include 47.5 mIoU in referring expression segmentation and 65.9 mIoU on Pascal-5i, indicating strong performance.
  • The framework utilizes adaptive prompt learning with cross-modal interaction to enhance image and text fusion for segmentation tasks.

Abstract

Open-vocabulary learning has revolutionized image segmentation, enabling the delineation of arbitrary categories from textual descriptions. While current methods often employ specialized architectures, OVAMTSeg presents a unified framework for Open-Vocabulary and Multitask Image Segmentation. Leveraging adaptive prompt learning, OVAMTSeg excels in capturing category-sensitive concepts, ensuring robustness across diverse multi-task scenarios. Text prompts effectively capture semantic and contextual features, while cross-attention and cross-modal interactions facilitate seamless fusion of image and text features. The framework incorporates a transformer-based decoder for dense prediction. Experimental results demonstrate OVAMTSeg's effectiveness, achieving a 47.5 mIoU in referring expression segmentation, 51.6 mIoU on Pascal-VOC with four unseen classes, 46.6 mIoU on Pascal-Context in zero-shot segmentation, 65.9 mIoU on Pascal-5i, and 35.7 mIoU on COCO-20i datasets for one-shot segmentation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pan et al. (2024) studied this question.

synapsesocial.com/papers/68e700f4b6db64358767b626https://doi.org/10.1145/3605098.3636192
Ask AI
Helpful
Bookmark
Share
View Full Paper