This work envisions a strong framework for improving video-based face recognition by addressing simultaneously the issues of pose and illumination changes. Since even exhaustive efforts in face recognition research have not eliminated the impact of dynamic environments in real-world applications on recognition accuracy caused by nonlinear facial distortions and lighting differences, this work is timely. To circumvent these challenges, a hybrid deep learning model is constructed that combines SENet and channel attention mechanisms for efficient spatial feature extraction and a Transformer network and cross-attention for temporal dependency modeling. The new system begins with face detection via SENResNet, feature refinement via Transformers, and stable tracking via the Regression Network-based Face Tracking (RNFT) model. This end-to-end system enables effective learning of invariant representations over diverse poses and lighting conditions. Testing on benchmark datasets shows substantial improvement in recognition accuracy and robustness, confirming the effectiveness of the proposed system in real-world applications of video surveillance and humancomputer interaction.
Gajanan S. Joshi (Fri,) studied this question.