We present MultiNet, a multi-branch pipeline that directly predicts the gaze vector or the point of gaze from multiple modalities derived from a single camera capture. In particular, we address the under-explored use of metric face-to-camera distance estimated from a conventional RGB web camera for appearance-based gaze estimation, in contrast to many distance-aware alternatives which require depth sensors, stereo rigs, or additional hardware. The proposed method employs a convolutional neural network (CNN) with attentional feature fusion and integrates cropped face images, left and right eye patches, head-pose information, and estimated camera-to-face distance. We also introduce NLGaze, a 141,000-image dataset collected for unconstrained and natural desktop settings, thereby emphasizing real-world applicability. In a person-specific real-time evaluation, MultiNet-Base reaches a root mean square error (RMSE) of 43.01 pixels at 15 frames per second (FPS) on a desktop with an off-the-shelf web-camera. We evaluate the proposed method against representative state-of-the-art solutions on benchmark datasets. MultiNet achieves angular errors of 2.94°, 7.47°, and 7.58° at best on the MPIIFaceGaze, Gaze360, and NLGaze datasets, respectively, demonstrating competitive or superior performance in appearance-based gaze estimation. The results of the ablation study show that the attentional feature fusion mechanism and the addition of the estimated distance input to the set of input modalities result in significant angular error reduction.
No takes yet. Share an insight, caveat, or question.
Ligostaev et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: