Key points are not available for this paper at this time.
Abstract With the rapid advancement of artificial intelligence, research on deep neural networks for estimating gaze from facial images captured by cameras has grown significantly. However, gaze estimation remains challenging due to the diversity of facial features, variability in head posture, partial occlusion, varying lighting conditions, eye image quality, complex eye movements, and a lack of large-scale, diverse data sets. Particularly, estimating a driver's gaze while driving a vehicle is difficult due to occlusion of the eyes when checking mirrors or turning the head to view blind spots and surroundings. To address these issues, we developed an appearance-based gaze estimation neural network based on attention, called GazeSymCAT (symmetric cross-attention transformer for gaze estimation). The proposed GazeSymCAT obtains contextual information that reflects the relationships among the extracted facial and eye image features through an encoder–decoder structure composed of self- and cross-attention layers of the transformer. This structure is followed by adding fully connected layers that estimate the gaze direction. The model proposed in this study was tested on the ETH-XGaze data set, a large-scale, state-of-the-art (SOTA) data set that includes extreme head postures, achieving SOTA performance. Additionally, it demonstrated SOTA-comparable performance on existing general gaze estimation-related data sets, MPIIFaceGaze and EYEDIAP.
Zhong et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: