PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 12, 2026IEEE Journal of Biomedical and Health Informatics1 citations

S 3 F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network

View Full Paper
MSMd. Saiful Bari SiddiquiMBMohammed Imamul Hassan Bhuiyan

Key Points

  • To enhance medical image classification by integrating spatial and spectral information using a dual-branch framework.
  • Proposed Spatial-Spectral Summarizer Fusion Network (S3F-Net) to learn spatial and spectral features simultaneously.
  • Utilized a deep spatial CNN in conjunction with a shallow spectral encoder, SpectraNet.
  • Applied the SpectralFilter layer for efficient processing of the Fourier spectrum through learnable filters.
  • Evaluated S3F-Net on four medical imaging datasets: HAM10000, BUSI, BRISC2025, and Chest X-Ray Pneumonia.
  • Achieved up to 5.13% accuracy improvement over a strong spatial-only baseline across all datasets.
  • Bilinear Fusion reached a competitive accuracy of 98.76% on the BRISC2025 dataset.
  • Concatenation Fusion achieved 93.11% accuracy on the Chest X-Ray Pneumonia dataset, outperforming deeper models.
  • Showed that reliance on spatial or spectral information adapts based on input pathology.

Abstract

Convolutional Neural Networks (CNNs) have become a cornerstone of medical image analysis due to their proficiency in learning hierarchical spatial features. However, this focus on a single domain is inefficient at capturing global, holistic patterns and fails to explicitly model an image's frequency-domain characteristics. To address these challenges, we propose the Spatial-Spectral Summarizer Fusion Network (S3F-Net), a dual-branch framework that learns from both spatial and spectral representations simultaneously. The S3F-Net performs a fusion of a deep spatial CNN with our proposed shallow spectral encoder, SpectraNet. SpectraNet features the proposed SpectralFilter layer, which leverages the Convolution Theorem by applying a bank of learnable filters directly to an image's full Fourier spectrum via a computation-efficient element-wise multiplication. This allows the SpectralFilter layer to attain a global receptive field instantaneously, with its output being distilled by a lightweight summarizer network. We evaluate S3F-Net across four diverse medical imaging datasets spanning different scales and modalities: HAM10000 (dermoscopy), BUSI (ultrasound), BRISC2025 (MRI), and Chest X-Ray Pneumonia (radiography), to validate its efficacy and generalizability, and reveal the task-dependent nature of the optimal fusion strategy. Our framework consistently and significantly outperforms its strong spatial-only baseline in all cases, with accuracy improvements of up to 5.13%. With a powerful Bilinear Fusion, S3F-Net achieves a state-of-the-art competitive accuracy of 98.76% on the BRISC2025 dataset. A simpler Concatenation Fusion performs better on the texture-dominant Chest X-Ray Pneumonia dataset, achieving 93.11% accuracy, surpassing many top-performing, much deeper models. Our explainability analysis also reveals that the S3F-Net learns to dynamically adjust its reliance on each branch based on the input pathology. These results verify that our dual-domain approach is a powerful and generalizable paradigm for medical image analysis. Our source codes and pre-trained model weights are publicly available on GitHub at: https://github.com/Saiful185/S3F-Net.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Siddiqui et al. (2026) studied this question.

synapsesocial.com/papers/69db361c4fe01fead37c46f8https://doi.org/10.1109/jbhi.2026.3682634
Ask AI
Helpful
Bookmark
Share
View Full Paper