WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Graphical User Interface (GUI) automation relies on accurate GUI grounding. However, obtaining large-scale, high-quality labeled data remains a key challenge, particularly in desktop environments like Windows Operating System (OS). Existing datasets primarily focus on structured web-based elements, leaving a gap in real-world GUI interaction data for non-web applications. To address this, we introduce a new framework that leverages LLMs to generate large-scale GUI grounding data, enabling automated and scalable labeling across diverse interfaces. To ensure high accuracy and reliability, we manually validated and refined 5,000 GUI coordinate-instruction pairs, creating WinSpot—the first benchmark specifically designed for GUI grounding tasks in Windows environments. WinSpot provides a high-quality dataset for training and evaluating visual GUI agents, establishing a foundation for future research in GUI automation across diverse and unstructured desktop environments1

Cite

CITATION STYLE

APA

Hui, Z., Li, Y., Zhao, D., Banbury, C., Chen, T., & Koishida, K. (2025). WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 1086–1096). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-short.85

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free