Randomized trial demonstrates improved performance in microexpression recognition, suggesting efficient modeling alternatives.
Microexpression recognition (MER) remains challenging because facial movements are short, subtle, and usually require computationally expensive temporal models. This paper presents a lightweight depthwise-separable Temporal Convolutional Network (DS-TCN) for MER on the CASME II dataset. The model uses a dual-stream design that combines grayscale facial-frame features with optical-flow motion features after face cropping, Eulerian video magnification, resizing, and SSIM-based key-frame selection. Standard temporal convolutions are replaced with depthwise separable temporal convolutions to reduce parameter count and floating-point operations while preserving sequence learning. Three-fold cross-validation produced a validation accuracy of 72.6% and a macro-F1 score of 67.58%. Compared with the standard TCN baseline, the proposed DS-TCN reduced parameters from 1.57 M to 0.219 M, FLOPs from 10.33 G to 1.50 G, and CPU inference time from 580 ms to 210 ms. The results show that lightweight temporal modeling can provide competitive MER performance with substantially lower computational cost.
No takes yet. Share an insight, caveat, or question.
Adeyemi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: