Randomized trial shows improved diagnostic reasoning in radiology using gaze data from experts, suggesting enhanced AI collaboration.
Large-scale vision-language models have shown promise in automating chest X-ray interpretation. However, their clinical utility remains limited, since most systems optimize for semantic information rather than emulating how experts visually examine and interpret medical images. As a result, current models often overlook critical findings, misrepresent anatomical context, or diverge from established diagnostic workflows. Radiologists, by contrast, follow structured protocols that sequentially assess anatomical regions, reducing missed findings and supporting reliable diagnostic reasoning. We therefore introduce Gaze-X, a vision-language model that leverages radiologists’ eye-tracking data as a behavioral prior for expert diagnostic reasoning. By incorporating gaze trajectories and fixation patterns into pretraining, Gaze-X learns to follow the spatial and temporal structure of radiologist attention. Using a curated dataset of over 30,000 key frames from five radiologists interpreting chest X-rays across diverse disease categories, we show that Gaze-X produces more accurate, interpretable, and expert-consistent outputs across a range of clinically relevant tasks. Unlike autonomous reporting systems, Gaze-X produces verifiable evidence artifacts, including inspection trajectories and finding-linked localized regions, enabling transparent and safe human-AI collaboration. This capability provides a practical route toward more trustworthy, explainable, and diagnostically robust AI for radiology and beyond.
No takes yet. Share an insight, caveat, or question.
Lee et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: