Abstract
Open-vocabulary image semantic segmentation (OVS) seeks to segment images into semantic regions across an open set of categories. Existing OVS methods commonly depend on foundational vision-language models and utilize similarity computation to tackle OVS tasks. However, these approaches are predominantly tailored to natural images and struggle with the unique characteristics of high-resolution remote sensing images, such as rapidly changing orientations and significant scale variations. To tackle this dilemma, we propose the first OVS framework specifically designed for high-resolution remote sensing imagery, introducing a rotation-aggregative similarity computation module to enhance segmentation across varying orientations and a multiscale feature integration strategy to generate scale-aware semantic masks. Additionally, we establish the first open-sourced OVS benchmark for remote sensing, comprising four public datasets. Experiments demonstrate that our framework effectively addresses orientation and scale challenges, achieving state-of-the-art performance.
Author supplied keywords
Cite
CITATION STYLE
Cao, Q., Chen, Y., Ma, C., & Yang, X. (2025). Open-Vocabulary High-Resolution Remote Sensing Image Semantic Segmentation. IEEE Transactions on Geoscience and Remote Sensing, 63. https://doi.org/10.1109/TGRS.2025.3559557
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.