Benchmarking Direct Preference Optimization for Medical Large Vision–Language Models

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Large Vision-Language Models (LVLMs) hold significant promise for medical applications, yet their deployment is often constrained by insufficient alignment and reliability. While Direct Preference Optimization (DPO) has emerged as a potent framework for refining model responses, its efficacy in high-stakes medical contexts remains underexplored, lacking the rigorous empirical groundwork necessary to guide future methodological advances. To bridge this gap, we present the first comprehensive examination of diverse DPO variants within the medical domain, evaluating nine distinct formulations across two medical LVLMs: LLaVA-Med and HuatuoGPT-Vision. Our results reveal several critical limitations: current DPO approaches often yield inconsistent gains over supervised fine-tuning, with their efficacy varying significantly across different tasks and backbones. Furthermore, they frequently fail to resolve fundamental visual misinterpretation errors. Building on these insights, we present a targeted preference construction strategy as a proof-of-concept that explicitly addresses visual misinterpretation errors frequently observed in existing DPO models. This design yields a 3.6% improvement over the strongest existing DPO baseline on visual question-answering tasks. To support future research, we release our complete framework, including all training data, model checkpoints, and our codebase at https://github.com/dmis-lab/med-vlm-dpo.

Cite

CITATION STYLE

APA

Kim, D., Lee, J., Yun, J., Koo, Y. H., Chen, Q., Kim, H., & Kang, J. (2026). Benchmarking Direct Preference Optimization for Medical Large Vision–Language Models. In 19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 (pp. 5052–5067). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2026.findings-eacl.267

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free