PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026Journal of King Saud University - Computer and Information Sciences2 citationsOpen Access

Hyperspectral image classification via Manhattan self-attention transformer and adaptive global-local channel attention

View Full Paper
ZMZhe MengPYPan YueFZFeng Zhao

Key Points

  • Overall accuracy of the proposed method reaches 99.11% on the Indian Pines dataset, exhibiting significant performance improvements.
  • The Manhattan self-attention transformer encoder enhances the modeling of long-range spatial dependencies across extracted features.
  • Wavelet-guided multi-scale spatial-spectral feature extraction facilitates the extraction of detailed local and global features.
  • Adaptive global-local channel attention allows more precise extraction of fine-grained spectral characteristics.

Abstract

Hyperspectral image (HSI) classification is a key task in the field of remote sensing with significant practical value. However, most convolutional neural network (CNN)-based methods primarily focus on capturing local spatial features while overlooking global context. In contrast, vision transformer (ViT)-based methods can effectively model long-range dependencies but fail to adequately capture detailed local spatial structures and lack explicit spatial priors in their self-attention mechanism. In this paper, we propose a novel CNN–ViT hybrid architecture for HSI classification, called the Manhattan self-attention transformer with adaptive global–local channel attention network (MTACANet). This network can simultaneously extract local and global features and incorporates an explicit, distance-based spatial prior into the attention mechanism. First, we introduce a wavelet-guided multi-scale spatial-spectral feature extraction (WMSFE) block to separately extract multi-scale spatial and spectral features, which begins with wavelet transform convolution (WTConv) featuring an expanded receptive field to strengthen feature representation capabilities. Second, we use a Manhattan self-attention transformer encoder (MSATE) to model long-range spatial dependencies across the extracted multi-scale spatial feature representations. The MSATE incorporates explicit spatial priors and adjusts attention weights based on inter-token distances, enabling each target token to focus more on nearby tokens while suppressing attention to distant ones. In addition, an adaptive global-local channel attention (AGLCA) module is employed to facilitate interactions between local and global channel information, allowing for more precise extraction of fine-grained spectral characteristics in HSIs. Experiments on the Indian Pines, GF-5 Yancheng, ZY1-02D Huanghekou, and WHU-Hi-HongHu datasets show that its overall accuracy reaches 99.11%, 99.28%, 99.09%, and 95.57%, respectively, exhibiting excellent performance. Our code is available at: https://github.com/zhe-meng/MTACANet.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Meng et al. (2026) studied this question.

synapsesocial.com/papers/69a75bc2c6e9836116a23ac2https://doi.org/10.1007/s44443-026-00484-1
Ask AI
Helpful
Bookmark
Share
View Full Paper