PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2026IET Biometrics1 citationsOpen Access

Multiscale Convolutional–Bidirectional LSTM Fusion of Spatiotemporal Attention for Human Activity Recognition

View Full Paper
YZYuyang ZhangCXChangcheng XiangYZYinrui Zhang

Key Points

  • The aim is to enhance human activity recognition through an innovative model that integrates multiple technologies.
  • Developed a model combining multiscale convolution, BiLSTM, and spatiotemporal attention.
  • Utilized filter sizes of 3, 5, 7, and 9 in a multiscale parallel convolutional structure.
  • Employed the Swish activation function to optimize feature extraction and address gradient issues.
  • Achieved an accuracy of 95.39% on the UCI-HAR public dataset.
  • Demonstrated excellent performance of 99.58% on a custom dataset.
  • Significantly improved accuracy through enhanced feature selection and noise resilience.

Abstract

Human activity recognition (HAR) is a core technology in fields such as smart healthcare and human–computer interaction, which aims to classify daily activities (e.g., walking and running) based on sensor data automatically. While existing approaches achieve high accuracy in controlled laboratory settings, they often perform poorly in real‐world applications due to limited multiscale feature extraction and poor modeling of interchannel sensor correlations. These shortcomings lead to high sensitivity to noise and irrelevant time segments. An innovative model integrating multiscale convolution, bidirectional long short‐term memory (BiLSTM), and a spatiotemporal attention mechanism is proposed to address these issues in this research. The model employs a multiscale parallel convolutional structure with filter sizes of 3, 5, 7, and 9 that enable it to capture both short‐term local dependencies and long‐term global patterns simultaneously. The introduction of a spatial–temporal dual attention mechanism dynamically focuses on key sensor channels and time segments, significantly improving the accuracy of feature selection. In addition, the Swish activation function is used to optimize the feature extraction process. Its smooth characteristics and self‐gating mechanism avoid the gradient vanishing problem of the traditional ReLU effectively. The experimental results show that the model achieves an accuracy of 95.39% on the UCI–HAR public dataset and an excellent performance of 99.58% on the custom dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/6a095c2c7880e6d24efe229chttps://doi.org/10.1049/bme2/8129974
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Long Short-Term Memory1997 · 101,723 citations
  2. 2Vision-based Human Activity Recognition Using Local Phase Quantization2024 · 5 citations
  3. 3Attention-Based Explainability Approaches in Healthcare Natural Language Processing2023 · 11 citations
  4. 4Deep Learning-Based Speed Bump Detection Model for Intelligent Vehicle System Using Raspberry Pi2020 · 97 citations
  5. 5TinyHAR: Benchmarking Human Activity Recognition Systems in Resource Constrained Devices2022 · 8 citations