PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 20250 citationsOpen Access

ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

View Full Paper
KLKaixin LiGuangdong University of TechnologyZMZiyang MengShanghai Zhaozhan Metal MaterialsHLHongzhan LinHong Kong Baptist University

Key Points

  • Achieving 48.1% accuracy with ScreenSeekeR in high-resolution professional environments demonstrates a significant improvement.
  • The benchmark utilizes authentic high-resolution images across 23 applications and five industries, showcasing real-world complexity.
  • Existing models struggle, with the best achieving only 18.9% on this challenging dataset, indicating areas needing improvement.
  • Strategically reducing the search area in GUI perception tasks shows promising results for enhancing accuracy in professional settings.

Abstract

Recent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays, smaller target sizes, and complex environments. In this paper, we introduce ScreenSpot-Pro, a new benchmark designed to rigorously evaluate the grounding capabilities of MLLMs in high-resolution professional settings. The benchmark comprises authentic high-resolution images from a variety of professional domains with expert annotations. It spans 23 applications across five industries and three operating systems. Existing GUI grounding models perform poorly on this dataset, with the best model achieving only 18.9%. Our experiments reveal that strategically reducing the search area enhances accuracy. Based on this insight, we propose ScreenSeekeR, a visual search method that utilizes the GUI knowledge of a strong planner to guide a cascaded search, achieving state-of-the-art performance with 48.1% without any additional training. We hope that our benchmark and findings will advance the development of GUI agents for professional applications. Code, data and leaderboard can be found at https://gui-agent.github.io/grounding-leaderboard.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68e24e59d6d66a53c2472eb6https://doi.org/10.48550/arxiv.2504.07981
Ask AI
Helpful
Bookmark
Share
View Full Paper