Understanding how humans interact with charts is crucial for designing effective visualization systems. While identifying a user’s task from gaze is fundamental, traditional methods rely on labor-intensive feature engineering, showing limited performance and high task-specificity. In this work, we introduce a framework for MLLM-based gaze-to-task inference, including an automatic few-shot sample generation process that creates structured demonstrations for in-context learning. We present the first systematic investigation into gaze-based task inference on charts, benchmarking five MLLMs and rigorously exploring gaze encoding and prompting strategies to establish optimal design principles. Our findings identify the heatmap representation as the optimal visual gaze encoding, and we demonstrate the necessity of Chain-of-Thought prompting. Notably, MLLMs can autonomously decode cognitive intent without manual AOI definitions, exceeding traditional baseline performance. This study offers actionable insights for integrating human gaze into MLLMs, guiding the design of future systems that adapt to a user’s analytical focus.
Nishiyasu et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: