LENS: Learning to Segment Anything with Unified Reinforced Reasoning

0Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing supervised fine-tuning methods typically ignore explicit chain-of-thought (CoT) reasoning at test time, which limits their ability to generalize to unseen prompts and domains. To address this issue, we introduce LENS, a scalable reinforcement-learning framework that jointly optimizes the reasoning process and segmentation in an end-to-end manner. We propose unified reinforcement-learning rewards that span sentence-, box-, and segment-level cues, encouraging the model to generate informative CoT rationales while refining mask quality. Using a publicly available 3-billion-parameter vision–language model, i.e., Qwen2.5-VL-3B-Instruct, LENS achieves an average cIoU of 81.2% on the RefCOCO, RefCOCO+, and RefCOCOg benchmarks, outperforming the strong fine-tuned method, i.e., GLaMM, by up to 5.6%. These results demonstrate that RL-driven CoT reasoning significantly enhances text-prompted segmentation and offers a practical path toward more generalizable Segment Anything models (SAM).

Cite

CITATION STYLE

APA

Zhu, L., Ouyang, B., Zhang, Y., Cheng, T., Hu, R., Shen, H., … Wang, X. (2026). LENS: Learning to Segment Anything with Unified Reinforced Reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, pp. 13952–13960). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v40i16.38405

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free