Abstract
Graphical User Interface (GUI) automation relies on accurate GUI grounding. However, obtaining large-scale, high-quality labeled data remains a key challenge, particularly in desktop environments like Windows Operating System (OS). Existing datasets primarily focus on structured web-based elements, leaving a gap in real-world GUI interaction data for non-web applications. To address this, we introduce a new framework that leverages LLMs to generate large-scale GUI grounding data, enabling automated and scalable labeling across diverse interfaces. To ensure high accuracy and reliability, we manually validated and refined 5,000 GUI coordinate-instruction pairs, creating WinSpot—the first benchmark specifically designed for GUI grounding tasks in Windows environments. WinSpot provides a high-quality dataset for training and evaluating visual GUI agents, establishing a foundation for future research in GUI automation across diverse and unstructured desktop environments1
Cite
CITATION STYLE
Hui, Z., Li, Y., Zhao, D., Banbury, C., Chen, T., & Koishida, K. (2025). WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 1086–1096). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-short.85
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.