Key result
Deep learning fusing ECG, voice, and facial expressions outperforms single modalities with ~85% stress detection accuracy.
Why the study?
Many studies use single-modality for stress detection and rarely combine stress-related information from multimodality.
Does a multimodality deep learning framework fusing ECG, voice, and facial expressions accurately detect acute mental stress?
Does a multimodality deep learning framework fusing ECG, voice, and facial expressions accurately detect acute mental stress?
Absolute Event Rate: 85.1% vs 83%
A multimodality deep learning framework fusing ECG, voice, and facial expressions can detect acute mental stress with 85.1% accuracy.
May support real-time multimodality stress monitoring; leaves open prospective validation and outcome studies before clinical use.
Mental stress is becoming increasingly widespread and gradually severe in modern society, threatening people's physical and mental health. To avoid the adverse effects of stress on people, it is imperative to detect stress in time. Many studies have demonstrated the effectiveness of using objective indicators to detect stress. Over the past few years, a growing number of researchers have been trying to use deep learning technology to detect stress. However, these works usually use single-modality for stress detection and rarely combine stress-related information from multimodality. In this paper, a real-time deep learning framework is proposed to fuse ECG, voice, and facial expressions for acute stress detection. The framework extracts the stress-related information of the corresponding input through ResNet50 and I3D with the temporal attention module (TAM), where TAM can highlight the distinguishing temporal representation for facial expressions about stress. The matrix eigenvector-based approach is then used to fuse the multimodality information about stress. To validate the effectiveness of the framework, a well-established psychological experiment, the Montreal imaging stress task (MIST), was applied in this work. We collected multimodality data from 20 participants during MIST. The results demonstrate that the framework can combine stress-related information from multimodality to achieve 85.1% accuracy in distinguishing acute stress. It can serve as a tool for computer-aided stress detection.
No takes yet. Share an insight, caveat, or question.
Zhang et al. (2022) studied Acute mental stress (n=20). Multimodality deep learning framework (ECG, voice, and facial expressions) vs. Single-modality detection methods was evaluated on Stress detection accuracy. A real-time deep learning framework fusing ECG, voice, and facial expressions achieved 85.1% accuracy in detecting acute mental stress, outperforming single-modality detection methods.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: