PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 3, 2026AI0 citationsOpen Access

A General Safety-Aware Hybrid Multimodal Architecture for Sign Language Understanding in Automated Vehicle Interaction

View Full Paper
SRSuresh RasappanFDFrancis Saviour DevarajASAhamed Nishath Syed

Key Points

  • This research aims to develop a hybrid multimodal architecture for effective sign language understanding in automated vehicle contexts, focusing on safety and reliability.
  • Proposes STCM-HVNet integrating visual encoder, pose encoder, temporal encoder, and safety decision module.
  • Implements multi-task predictions for sign category, interaction intent, and urgency level.
  • Conducts experiments on RGBArS image benchmark and Arabic sign-language video benchmark with detailed accuracy metrics.
  • On the RGBArS image dataset, achieved Top-1 accuracy of 45.38% and Macro-F1 of 0.4479.
  • On Arabic sign-language video benchmark, the BiLSTM temporal encoder reached Top-1 accuracy of 93.15% and Macro-F1 of 0.9383.
  • Safety decision layer offers a balance between prediction coverage and reliability as shown by Monte Carlo dropout analysis.

Abstract

Sign language understanding for automated vehicles sits at the intersection of accessibility, intelligent transportation, and safety-critical human–machine interaction. The existing sign-language recognition systems are largely confined to controlled environments, limiting their utility in mobility scenarios characterized by lighting variation, motion blur, and partial occlusion. This paper proposes STCM-HVNet, a safety-aware hybrid multimodal architecture integrating four components: a spatial visual encoder, a MediaPipe-based pose encoder, a bidirectional LSTM temporal encoder, and a context-aware fusion and safety decision module. The architecture is formulated as a multi-task system that jointly predicts sign category, interaction intent, and urgency level, and incorporates confidence-aware rejection and fail-safe action mapping. Experiments are conducted on two Arabic sign-language resources. On the RGBArS image benchmark (31 classes, 7856 images), the proposed pipeline achieves a Top-1 accuracy of 45.38%, Top-3 accuracy of 75.15%, and Macro-F1 of 0.4479, outperforming LinearECOC, kNN-5, and Bagged Trees baselines. On the Arabic sign-language video benchmark (12 classes, 479 clips), the BiLSTM temporal encoder achieves a Top-1 accuracy of 93.15% and Macro-F1 of 0.9383, outperforming frame-aggregation (87.67%) and CNN-LSTM (89.04%) baselines. Ablation results confirm complementary contributions from the visual and pose branches. A safety-threshold analysis and a Monte Carlo dropout comparison demonstrate that the proposed safety decision/gating layer provides a controllable trade-off between prediction coverage and reliability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rasappan et al. (2026) studied this question.

synapsesocial.com/papers/6a1fc616dee9eb8c0dce748dhttps://doi.org/10.3390/ai7060200
Ask AI
Helpful
Bookmark
Share
View Full Paper