Abstract
Open-vocabulary semantic segmentation (OVSS) aims to segment objects without being constrained by a predefined set of categories. Recent advancements in OVSS have been driven by CLIP, a powerful vision-language model that enables segmentation through textual and visual feature matching. However, while visual features within a class exhibit significant diversity, their corresponding text features remain relatively limited in distribution. This discrepancy weakens the effectiveness of vision-language matching for OVSS. To address these challenges, we propose OV-EVA (Open-Vocabulary Semantic Segmentation with Enhanced Vision-Language Alignment), a novel training-free OVSS framework that leverages multi-modal large language models (MLLMs) to improve visual-language alignment and enhance segmentation robustness. Our method introduces an iterative vocabulary expansion strategy, where MLLMs generate a diverse and scene-relevant vocabulary set through a two-stage querying process. To ensure accurate segmentation, we incorporate a mask-guided score refinement mechanism, which enhances vocabulary terms closely aligned with the target mask while mitigating overly dominant terms across the entire image. Additionally, we introduce a top-N -based target class mapping strategy to improve the alignment between the generated vocabulary set and target class labels. The proposed OV-EVA achieves state-of-the-art performance across multiple benchmarks when using GPT-4o as the underlying MLLM. Furthermore, our approach demonstrates strong adaptability across different MLLMs, achieving competitive results with LLaVA and Janus-Pro. These findings suggest that leveraging generative reasoning from MLLMs offers a scalable pathway for robust zero-shot perception without task-specific training. Future work will focus on optimizing computational efficiency to facilitate real-time deployment in practical applications. The source code is available at https://github.com/c-jang/ov-eva.
Author supplied keywords
Cite
CITATION STYLE
Jang, C., Youn, S. J., Kang, I., Kwon, J., Ji, D., Choi, J. W., & Cho, N. I. (2026). Robust Vision-Language Alignment Using Multi-Modal Large Language Models for Open-Vocabulary Semantic Segmentation. IEEE Access, 14, 24410–24431. https://doi.org/10.1109/ACCESS.2026.3663647
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.