PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 17, 2025Applied Sciences0 citationsOpen Access

UNETR with Voxel-Focused Attention: Efficient 3D Medical Image Segmentation with Linear-Complexity Transformers++

View Full Paper
SNSithembiso NtanziSVSerestina Viriri

Key Points

  • The voxel-focused attention mechanism achieves linear complexity, enhancing efficiency in 3D medical image segmentation.
  • The improved UNETR++ model reduces parameters by 50%, reaching 21.42 M while maintaining a competitive Dice score of 86.72%.
  • The Efficient Paired Attention block incorporates both spatial and channel attention, optimizing model performance.
  • This research addresses the challenge of information loss associated with dimensionality reduction in transformer models.

Abstract

There have been significant breakthroughs in developing models for segmenting 3D medical images, with many promising results attributed to the incorporation of Vision Transformers (ViT). However, the fundamental mechanism of transformers, known as self-attention, has quadratic complexity, which significantly increases computational requirements, especially in the case of 3D medical images. In this paper, we investigate the UNETR++ model and propose a voxel-focused attention mechanism inspired by TransNeXt pixel-focused attention. The core component of UNETR++ is the Efficient Paired Attention (EPA) block, which learns from two interdependent branches: spatial and channel attention. For spatial attention, we incorporated the voxel-focused attention mechanism, which has linear complexity with respect to input sequence length, rather than projecting the keys and values into lower dimensions. The deficiency of UNETR++ lies in its reliance on dimensionality reduction for spatial attention, which reduces efficiency but risks information loss. Our contribution is to replace this with a voxel-focused attention design that achieves linear complexity without low-dimensional projection, thereby reducing parameters while preserving representational power. This effectively reduces the model’s parameter count while maintaining competitive performance and inference speed. On the Synapse dataset, the enhanced UNETR++ model contains 21.42 M parameters, a 50% reduction from the original 42.96 M, while achieving a competitive Dice score of 86.72%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ntanzi et al. (2025) studied this question.

synapsesocial.com/papers/68f199c5de32064e504dcd5bhttps://doi.org/10.3390/app152011034
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1UNETR++: Delving Into Efficient and Accurate 3D Medical Image Segmentation2024 · 425 citations
  2. 2Attention-Enhanced Bimodal 3D Medical Image Segmentation with Two-Stage Learning2026
  3. 3TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers2024 · 1,209 citations
  4. 4TokenUNet: A new case for transformers integration in efficient and interpretable 3D UNets for brain imaging segmentation2026
  5. 5Dual Stream Fusion U-Net Transformers for 3D Medical Image Segmentation2024