A unified framework integrates visual and wearable sensors to enhance human motion monitoring, indicating potential improvements in biomedical applications.
This study proposes a unified multimodal temporal motion state perception framework for optical imaging-oriented biomedical applications, integrating visual skeleton sequences, inertial measurement unit (IMU) signals, and surface electromyography (EMG) signals. The framework utilizes modality-specific encoders and a cross-modal temporal alignment attention mechanism to explicitly model temporal offsets from heterogeneous sensing streams. A multimodal temporal Transformer backbone is introduced to capture long-range motion dependencies and cross-modal interactions, while an uncertainty-aware fusion module dynamically allocates weights based on modality confidence. Experimental results demonstrate that the proposed approach achieves an accuracy of 94.37%, an F1-score of 93.95%, and a mean average precision of 96.02%, outperforming mainstream baseline models. Robustness evaluations further confirm stable performance under visual occlusion and sensor noise. These results indicate that the framework provides a highly accurate and robust solution for rehabilitation assessment, sports training monitoring, and wearable intelligent interaction systems.
No takes yet. Share an insight, caveat, or question.
Chen et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: