Robust Vision-Language Alignment Using Multi-Modal Large Language Models for Open-Vocabulary Semantic Segmentation

0Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Open-vocabulary semantic segmentation (OVSS) aims to segment objects without being constrained by a predefined set of categories. Recent advancements in OVSS have been driven by CLIP, a powerful vision-language model that enables segmentation through textual and visual feature matching. However, while visual features within a class exhibit significant diversity, their corresponding text features remain relatively limited in distribution. This discrepancy weakens the effectiveness of vision-language matching for OVSS. To address these challenges, we propose OV-EVA (Open-Vocabulary Semantic Segmentation with Enhanced Vision-Language Alignment), a novel training-free OVSS framework that leverages multi-modal large language models (MLLMs) to improve visual-language alignment and enhance segmentation robustness. Our method introduces an iterative vocabulary expansion strategy, where MLLMs generate a diverse and scene-relevant vocabulary set through a two-stage querying process. To ensure accurate segmentation, we incorporate a mask-guided score refinement mechanism, which enhances vocabulary terms closely aligned with the target mask while mitigating overly dominant terms across the entire image. Additionally, we introduce a top-N -based target class mapping strategy to improve the alignment between the generated vocabulary set and target class labels. The proposed OV-EVA achieves state-of-the-art performance across multiple benchmarks when using GPT-4o as the underlying MLLM. Furthermore, our approach demonstrates strong adaptability across different MLLMs, achieving competitive results with LLaVA and Janus-Pro. These findings suggest that leveraging generative reasoning from MLLMs offers a scalable pathway for robust zero-shot perception without task-specific training. Future work will focus on optimizing computational efficiency to facilitate real-time deployment in practical applications. The source code is available at https://github.com/c-jang/ov-eva.

Cite

CITATION STYLE

APA

Jang, C., Youn, S. J., Kang, I., Kwon, J., Ji, D., Choi, J. W., & Cho, N. I. (2026). Robust Vision-Language Alignment Using Multi-Modal Large Language Models for Open-Vocabulary Semantic Segmentation. IEEE Access, 14, 24410–24431. https://doi.org/10.1109/ACCESS.2026.3663647

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free