PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 26, 2026Scientific Reports2 citationsOpen Access

Real-time underwater object detection via frequency-domain dynamics and spatially enhanced feature modulation

SCShaobin CaiAZAocheng Zhu

Key Points

  • To develop a lightweight framework for real-time underwater object detection that balances accuracy and latency.
  • Designed FasterFDBlock backbone integrating Partial Convolution and Frequency-domain Dynamic Convolution.
  • Introduced AIFI-SEFN encoder for enhancing small object detail through a Spatially Enhanced Feed-Forward Network.
  • Applied Multi-scale Feature Modulation (MFM) module to weigh deep semantic and shallow detailed features.
  • Achieved a mean Average Precision (mAP) of 72.1%, outperforming the baseline by 1.7%.
  • Reduced parameters and GFLOPs by 27.1% and 24.6%, respectively.
  • Accomplished an inference speed of 72.6 FPS.

Abstract

Underwater object detection is crucial for marine engineering yet challenged by complex optical properties causing blurring, low contrast, and texture loss. Furthermore, deploying deep learning models on resource-constrained platforms necessitates balancing accuracy with latency. We propose a novel lightweight framework based on the Real-Time Detection Transformer (RT-DETR). First, we design the FasterFDBlock backbone, integrating Partial Convolution with Frequency-domain Dynamic Convolution. This utilizes frequency band modulation to adaptively suppress high-frequency noise and enhance edge details, optimizing feature extraction with reduced redundancy. Second, to mitigate small object detail loss, we introduce the AIFI-SEFN encoder, incorporating a Spatially Enhanced Feed-Forward Network to synthesize global semantic context with local spatial data. Third, a Multi-scale Feature Modulation (MFM) module is applied to dynamically weight deep semantic and shallow detailed features, bolstering robustness against scale variations and background interference. Experimental results on the UTDAC2020 dataset show our method achieves a mean Average Precision (mAP) of 72.1%, outperforming the baseline by 1.7%. Crucially, parameters and Floating Point Operations (GFLOPs) are reduced by 27.1% and 24.6%, respectively. With an inference speed of 72.6 FPS, this model offers a highly efficient solution for real-time underwater perception.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cai et al. (2026) studied this question.

synapsesocial.com/papers/69c4ccc9fdc3bde448918626https://doi.org/10.1038/s41598-026-44628-9
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Underwater object detection and datasets: a survey2024 · 82 citations
  2. 2Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization2017 · 22,825 citations
  3. 3Gradient-based learning applied to document recognition1998 · 59,548 citations
  4. 4Deep Residual Learning for Image Recognition2016 · 228,344 citations
  5. 5Color Balance and Fusion for Underwater Image Enhancement2017 · 1,342 citations