The growing demand for seamless and intuitive interaction in virtual and augmented reality environments has significantly accelerated advancements in Facial Expression Recognition (FER). While convolutional neural networks (CNNs) have demonstrated effectiveness in FER, challenges such as pose variation, computational complexity, and scalability continue to hinder real-world deployment. In this study, we propose a novel high performance computing (HPC) based pose-aware FER framework (HPCFER), termed that integrates Contrastive Learning, Swin Transformers through massive parallel computing to enhance both recognition accuracy and computational efficiency. The proposed dual branch hybrid architecture combines local feature extraction via pre-trained CNNs with global context modeling through Swin Transformers. To ensure pose invariance and robust generalization, pose-aware contrastive learning aligns expression embeddings across diverse head orientations, while a Stacked Neural Network Ensemble (SNNE) dynamically fuses multiple transformer-based classifiers using confidence and pose-aware weighting. Spatial Transformer Networks (STNs) further improve facial alignment under extreme pose variations. To address the intensive computational demands of this architecture, we integrate MPI for distributed cluster-level training, OpenMP for optimized multi-threaded CPU execution, and CUDA based GPU acceleration, enabling scalable, massively parallel learning and real time inference. Experimental evaluations on the RAF-DB dataset demonstrate that our proposed scheme achieves state-of-the-art accuracy of 97. 2% while significantly reducing training and inference time, highlighting its strong potential for real world, high throughput applications in immersive virtual and augmented reality systems.
Alghamdi et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: