PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 2026Information0 citationsOpen Access

Dual-View Sign Language Recognition via Front-View Guided Feature Fusion for Automatic Sign Language Training

SJSiyuan JingGYGaorong Yan

Key Points

  • The aim is to enhance automatic sign language training through improved word-level sign language recognition.
  • Developed an algorithm for word-level sign language recognition on the NationalCSL-DP dataset.
  • Performed frame-level alignment of dual-view sign videos.
  • Used a two-stage deep neural network to extract spatiotemporal features of signers.
  • Introduced front-view guided early fusion for feature integration.
  • Conducted extensive experiments to evaluate algorithm performance.
  • The proposed algorithm outperformed existing dual-view recognition algorithms.
  • Achieved Top-1 accuracy that is 10.29% higher than MViT and 11.38% higher than CNN + Transformer.
  • Demonstrated significant improvements in degrading effects of hand occlusion and limited datasets.

Abstract

The foundation of an automatic sign language training (ASLT) system lies in word-level sign language recognition (WSLR), which refers to the translation of captured sign language signals into sign words. However, two key issues need to be addressed in this field: (1) the number of sign words in all public sign language datasets is too small, and the words do not match real-world scenarios, and (2) only single-view sign videos are typically provided, which makes solving the problem of hand occlusion difficult. In this work, we design an efficient algorithm for WSLR which is trained on our recently released NationalCSL-DP dataset. The algorithm first performs frame-level alignment of dual-view sign videos. A two-stage deep neural network is then employed to extract the spatiotemporal features of the signers, including hand motions and body gestures. Furthermore, a front-view guided early fusion (FvGEF) strategy is proposed for effective fusion of features from different views. Extensive experiments were carried out to evaluate the algorithm. The results show that the proposed algorithm significantly outperformed existing dual-view sign language recognition algorithms. Compared with several state-of-the-art methods, the proposed algorithm achieves Top-1 accuracy on the NationalCSL6707 dataset that is 10.29 and 11.38 higher than MViT and CNN + Transformer, respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jing et al. (2026) studied this question.

synapsesocial.com/papers/698828620fc35cd7a8847e07https://doi.org/10.3390/info17020158
Ask AI
Helpful
Bookmark
Share
View Full Paper