PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 10, 20262 citationsOpen Access

Lightweight Multi-Scale Framework for Human Pose and Action Classification

ASAlireza SaberMHM.M. HosseiniAFAmirreza Fateh

Key Points

  • The study aims to improve human pose and action classification using a lightweight modular attention-based architecture.
  • Developed a modular attention-based architecture leveraging a Swin Transformer backbone.
  • Integrated Spatial Attention, Context-Aware Channel Attention, and Dual Weighted Cross Attention modules.
  • Employed explainable AI techniques for improved model reliability and interpretability.
  • Evaluated the model on Yoga-82 and Stanford 40 Actions datasets.
  • Achieved accuracies of 90.40% in 6-class and 87.44% in 20-class Yoga-82 configurations.
  • Attained an accuracy of 94.28% for the Stanford 40 Actions dataset.
  • Outperformed state-of-the-art models in accuracy, precision, recall, F1-score, and mean average precision with only 0.79 million parameters.

Abstract

Human pose classification, along with related tasks such as action recognition, is a crucial area in deep learning due to its wide range of applications in assisting human activities. Despite significant progress, it remains a challenging problem because of high inter-class similarity, dataset noise, and the large variability in human poses. In this paper, we propose a lightweight yet highly effective modular attention-based architecture for human pose classification, built upon a Swin Transformer backbone for robust multi-scale feature extraction. The proposed design integrates the Spatial Attention module, the Context-Aware Channel Attention Module, and a novel Dual Weighted Cross Attention module, enabling effective fusion of spatial and channel-wise cues. Additionally, explainable AI techniques are employed to improve the reliability and interpretability of the model. We train and evaluate our approach on two distinct datasets: Yoga-82 (in both main-class and subclass configurations) and Stanford 40 Actions. Experimental results show that our model outperforms state-of-the-art baselines across accuracy, precision, recall, F1-score, and mean average precision, while maintaining an extremely low parameter count of only 0.79 million. Specifically, our method achieves accuracies of 90.40% and 87.44% for the 6-class and 20-class Yoga-82 configurations, respectively, and 94.28% for the Stanford 40 Actions dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Saber et al. (2026) studied this question.

synapsesocial.com/papers/698acac07c832249c30ba0afhttps://doi.org/10.3390/s26041102
Ask AI
Helpful
Bookmark
Share
View Full Paper