ABSTRACT In Industrial IoT (IIoT) and Industry 5.0 settings, accurate understanding of human operations is essential for safety, ergonomics, and human–machine collaboration. Existing systems often rely on single‐modality sensing and cannot capture complex multimodal spatiotemporal patterns. This study proposes MMSTAHN, a Multimodal Spatiotemporal Attention Hybrid Network that integrates dynamic attention, a multi‐scale Conv1D–BiLSTM–Transformer encoder, domain‐adversarial adaptation, and an interpretable semantic decoder. MMSTAHN achieves an F1 score of 88.9%, outperforming Transformer models by 10.6%. In IIoT transfer scenarios, accuracy improves by 22.1% with < 40 ms latency, enabling real‐time edge deployment. The model identifies key operational phases and modality contributions, providing a robust, interpretable framework for safety monitoring, skill evaluation, and collaborative intelligence in Industry 5.0 environments.
Ma et al. (Thu,) studied this question.