Key points are not available for this paper at this time.
Saliency prediction models are typically trained on natural images, focusing on features such as shape and color. However, predicting saliency in images with text is challenging because the human brain processes text differently than it processes visual objects. To address this research gap, we fine-tuned a saliency model to improve the accuracy of images containing text, specifically, movie posters. Our fine-tuned model—based on GSGNet and TranSalNet—outperformed the original models in predicting the saliency map for movie posters. The experimental results indicate that text elements exhibit patterns that can be learned for better saliency prediction. • Text has semantic aspects that impact the saliency map. • Fine-tuning can improve the model’s performance in predicting images containing text. • There are patterns when someone views a text, and these can be learned by the model.
Nugraha et al. (Sun,) studied this question.