PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026IEEE Transactions on Medical Imaging6 citations

Text-Image Co-Alignment for Weakly Supervised Polyp Segmentation

View Full Paper
WHWenhui HuangZPZhen PanXWXiaoyan Wang

Key Points

  • To develop a framework for polyp segmentation that integrates text-driven weak supervision to enhance accuracy and reduce reliance on annotations.
  • Proposed Text-Image Co-Alignment (TICoA) for segmenting polyps using structured clinical descriptions as supervision.
  • Utilized contrastive learning to connect specific phrases with their related image areas.
  • Implemented a State-Space Model for efficiently modeling dependencies and a Fusion module for cross-modal interaction.
  • TICoA demonstrated competitive performance against leading weakly supervised segmentation methods.
  • Achieved improved semantic grounding and accuracy in polyp segmentation tasks compared to existing approaches.

Abstract

Fully supervised polyp segmentation relies on costly pixel-level annotations. Although semi- and weakly supervised methods reduce annotation requirements, they still depend on partial mask supervision. Text-supervised segmentation is a promising alternative; however, for polyps, the key challenge is to ground instance-specific phrases to the correct lesion region under cluttered backgrounds and large appearance variations. Existing approaches often rely on coarse text-image alignment, limiting precise region-level semantic correspondence. In this paper, we propose Text-Image Co-Alignment (TICoA), a text-supervised framework for polyp segmentation. TICoA leverages large language models (LLMs)-generated structured clinical descriptions as weak supervision and formulates segmentation as a fine-grained phrase-region coalignment problem. Through contrastive learning, TICoA explicitly associates query phrases with corresponding image regions to achieve robust semantic grounding under weak supervision. Architecturally, we adopt a State-Space Model (Mamba) to efficiently model long-range dependencies with linear computational complexity. To support effective cross-modal interaction, we further design a dedicated Mamba Fusion module with a Bi-Dimension Fusion (BiDF) strategy, which progressively propagates information along spatial and channel dimensions. Experiments on polyp datasets, with additional validation on skin lesion segmentation, demonstrate that TICoA is competitive with state-of-the-art weakly supervised methods. Our code and data are available at https://github.com/silentyuchen/TICoA.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2026) studied this question.

synapsesocial.com/papers/69be34f26e48c4981c6731f6https://doi.org/10.1109/tmi.2026.3674592
Ask AI
Helpful
Bookmark
Share
View Full Paper