Psychological stress is a critical factor affecting health, necessitating accurate assessment in fields such as healthcare and human-computer interaction. Traditional stress detection methods, involving direct contact or self-report scales, can compromise comfort and privacy. To address these limitations, this paper proposes a machine vision-based psychological stress perception strategy for human-computer interaction scenarios, fusing facial behavioral and physiological data to identify psychological stress states (stress and non-stress) with low interference and high accuracy. A deep learning network framework, combining siamese networks and two-stage learning, is introduced. This method employs multidimensional data, including remote photoplethysmography (rPPG), gaze angle gradient, and head posture gradient, all of which can be obtained via a camera without direct contact, ensuring privacy and comfort. To validate the framework's performance, stratified ten-fold cross-validation is conducted on the public UBFC-Phys dataset. The results demonstrate the proposed framework surpasses existing comparative algorithms, achieving a detection accuracy of 95.39% on the test set. Furthermore, the fusion of facial behavior and physiological data enhances classification performance, with an accuracy increase of 4% over facial behavior data alone and 2% over physiological data alone. This study enriches the theoretical framework of psychological stress assessment and provides technical support for intelligent stress monitoring.
Song et al. (2026) studied this question.