PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 19, 2026IEEE Transactions on Cybernetics3 citationsOpen Access

Learning Conditional Diffusion Transformer for Salient Object Detection in Optical Remote Sensing Images

View Full Paper
CZChao ZengLZLi ZhangSKSam Kwong

Key Points

  • The aim is to improve salient object detection in optical remote-sensing images using a new transformer architecture.
  • Developed a conditional diffusion transformer network (CDTNet) for detecting salient objects.
  • Implemented a progressive cross-stage fusion (PCSF) module for integrating multiscale features.
  • Applied a patch strategy (PS) for fine-grained feature aggregation.
  • Enhanced features from the backbone network using an encoder feature enhancement (EFE) module.
  • Experimental results show CDTNet outperforms existing state-of-the-art methods.
  • The approach significantly improves the accuracy of saliency prediction in remote sensing images.

Abstract

In recent years, the task of detecting salient objects in optical remote-sensing images has posed a significant and formidable challenge. The existing approaches heavily rely on a limited amount of label saliency masks and usually utilize convolutional neural networks (CNNs) for feature decoding. In this article, we introduce the conditional diffusion transformer network (CDTNet), a novel architecture meticulously designed to learn contextualized and diffusion-guided features for optical remote sensing image salient object detection (ORSI SOD). Our work presents a Transformer-based progressive cross-stage fusion (PCSF) module. This module serves as the decoding unit for saliency prediction, enabling the seamless integration of multiscale features from different stages of the network. Through this fusion, the model can better understand the inner structure of the image and enhance the accuracy of saliency prediction. Moreover, we develop a patch strategy (PS). This strategy is dedicated to fine-grained feature aggregation, allowing the network to focus on detailed information within individual feature patches and thus making better use of transformer layers. In addition, the encoder feature enhancement (EFE) module is applied to enhance the extracted features from the backbone network by utilizing spatial and channel attention. We conduct comprehensive experiments on various benchmark datasets and evaluation metrics. The experimental results unequivocally demonstrate the superiority of the proposed CDTNet over the comparison SOTA methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zeng et al. (2026) studied this question.

synapsesocial.com/papers/69bb9212496e729e6297f463https://doi.org/10.1109/tcyb.2026.3667145
Ask AI
Helpful
Bookmark
Share
View Full Paper