Visual saliency modeling has achieved high predictive performance in natural image domains, yet its generalization to abstract art remains limited by the lack of explicit semantic structure and the scarcity of eye-tracking data. In such semantically ambiguous contexts, understanding the underlying drivers of attention is as critical as predictive accuracy. This paper presents an interpretable, ’white-box’ saliency framework tailored to abstract art, which constructs predictions through a weighted combination of 35 modular heuristics grounded in perceptual psychology and art theory, including contrast, grouping, isolation and symmetry. Heuristic weights are optimized via a genetic algorithm and refined by a context-aware modulation mechanism that adapts to image-level visual features. Evaluation against eye-tracking data from 40 abstract paintings demonstrates that the model with the expanded activation variant produces stable, meaningful predictions while achieving a competitive KL-divergence score (1.11 ± 0.55), which is comparable to the SalGAN baseline (1.11 ± 0.53). Analysis of the optimized weights reveals strong contributions from contrast, texture, and grouping mechanisms, while nearly half of the heuristics, including most horizontal symmetry heuristics are systematically pruned by the model. Moreover, context-aware modulation reveals that these weights are not static but shift dynamically based on image-level features such as edge density and intensity variation. By prioritizing transparency over raw predictive performance, this study demonstrates that explainable saliency models can function as robust investigative tools for decoding the principles of human visual perception in data-scarce domains.
Vaičekauskas et al. (2026) studied this question.