PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 13, 2026Scientific Reports0 citationsOpen Access

Skeleton motion topology-masked prediction and contrastive learning for self-supervised human action recognition

YHYan Keung HuiFLFengyu LiXHXiuhua Hu

Key Points

  • This research aims to improve human action recognition using a hybrid self-supervised framework that accounts for joint dependencies and data augmentation.
  • Proposes a topology-masked motion modeling approach to encode motion dynamics and skeletal topology.
  • Employs a multi-stage hybrid augmentation strategy combining conventional and extreme methods.
  • Introduces a trajectory-guided feature dropping module to selectively discard non-critical features.
  • Utilizes large-scale unlabeled skeleton data for self-supervised learning.
  • The model achieves superior performance in recognizing actions, especially in occluded and complex environments.
  • Significant improvement in action recognition accuracy under low-supervision conditions.
  • Effectively reduces reliance on annotated datasets and mitigates visual interference.

Abstract

To address the limitations in data augmentation and neglect of joint dependencies in self-supervised human action recognition, this paper proposes a hybrid framework that integrates topology-masked motion modeling with contrastive learning. The proposed motion topology-masking technique jointly encodes skeletal topology and motion dynamics, preventing the model from over-focusing on temporally salient regions of prominent motions. We employ a multi-stage hybrid augmentation strategy, combining conventional and extreme augmentation methods to generate diverse, enriched positive pairs for contrastive learning. Additionally, we introduce a trajectory-guided feature dropping module, which selectively discards critical features based on trajectory attention maps, preventing the model from avoiding excessive focus on local joint trajectories. This approach effectively leverages large-scale unlabeled skeleton data through self-supervised learning, significantly reducing reliance on costly annotated datasets. Extensive experiments on NTU-60, NTU-120, and PKU-MMD demonstrate that the proposed model achieves superior performance in both occluded scenarios under complex environments and low-supervision conditions. It effectively mitigates visual interference and annotation scarcity while substantially improving action recognition accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hui et al. (2026) studied this question.

synapsesocial.com/papers/698ebeb185a1ff6a930160dchttps://doi.org/10.1038/s41598-026-39330-9
Ask AI
Helpful
Bookmark
Share
View Full Paper