Randomized trial demonstrates effective micro-expression recognition using advanced dual-stream network, suggesting improved accuracy in affective computing.
Micro-expression recognition (MER) remains a formidable challenge in affective computing due to the subtle, localized, and fleeting nature of facial muscle actuations. Conventional spatial-temporal networks are easily overwhelmed by static facial topologies, leading to feature representations that are heavily biased toward identity-specific noise. To address this, we propose the Motion-Anchored Cross-Modal Fusion Network (MACFN), a novel dual-stream ViT architecture that explicitly decouples and synergizes spatial appearance and optical flow dynamics. Specifically, we introduce a motion-anchored spatial attention module, which translates latent motion features into a sparse spatial probability mask. It acts as an enhancement gate, forcing the texture stream to bypass static backgrounds and attend to genuine ME-related regions. Furthermore, we design a cross-modal bilinear fusion module to capture the second-order interactions across modalities, mapping the coupled features into a discriminative semantic manifold. Extensive experiments conducted on the CASME II, SAMM, and SMIC databases under the rigorous leave-one-subject-out composite database evaluation protocol demonstrate that MACFN is effective and achieves competitive performance compared to several recent methods.
No takes yet. Share an insight, caveat, or question.
Li et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: