Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Dain, Lee, Jiwoo, Yun, Jaehoon, Koo, Yong Hoe, Chen, Qingyu, Kim, Hyunjae, Kang, Jaewoo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914279876198400
author Kim, Dain
Lee, Jiwoo
Yun, Jaehoon
Koo, Yong Hoe
Chen, Qingyu
Kim, Hyunjae
Kang, Jaewoo
author_facet Kim, Dain
Lee, Jiwoo
Yun, Jaehoon
Koo, Yong Hoe
Chen, Qingyu
Kim, Hyunjae
Kang, Jaewoo
contents Large Vision-Language Models (LVLMs) hold significant promise for medical applications, yet their deployment is often constrained by insufficient alignment and reliability. While Direct Preference Optimization (DPO) has emerged as a potent framework for refining model responses, its efficacy in high-stakes medical contexts remains underexplored, lacking the rigorous empirical groundwork necessary to guide future methodological advances. To bridge this gap, we present the first comprehensive examination of diverse DPO variants within the medical domain, evaluating nine distinct formulations across two medical LVLMs: LLaVA-Med and HuatuoGPT-Vision. Our results reveal several critical limitations: current DPO approaches often yield inconsistent gains over supervised fine-tuning, with their efficacy varying significantly across different tasks and backbones. Furthermore, they frequently fail to resolve fundamental visual misinterpretation errors. Building on these insights, we present a targeted preference construction strategy as a proof-of-concept that explicitly addresses visual misinterpretation errors frequently observed in existing DPO models. This design yields a 3.6% improvement over the strongest existing DPO baseline on visual question-answering tasks. To support future research, we release our complete framework, including all training data, model checkpoints, and our codebase at https://github.com/dmis-lab/med-vlm-dpo.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17918
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
Kim, Dain
Lee, Jiwoo
Yun, Jaehoon
Koo, Yong Hoe
Chen, Qingyu
Kim, Hyunjae
Kang, Jaewoo
Computer Vision and Pattern Recognition
Computation and Language
Large Vision-Language Models (LVLMs) hold significant promise for medical applications, yet their deployment is often constrained by insufficient alignment and reliability. While Direct Preference Optimization (DPO) has emerged as a potent framework for refining model responses, its efficacy in high-stakes medical contexts remains underexplored, lacking the rigorous empirical groundwork necessary to guide future methodological advances. To bridge this gap, we present the first comprehensive examination of diverse DPO variants within the medical domain, evaluating nine distinct formulations across two medical LVLMs: LLaVA-Med and HuatuoGPT-Vision. Our results reveal several critical limitations: current DPO approaches often yield inconsistent gains over supervised fine-tuning, with their efficacy varying significantly across different tasks and backbones. Furthermore, they frequently fail to resolve fundamental visual misinterpretation errors. Building on these insights, we present a targeted preference construction strategy as a proof-of-concept that explicitly addresses visual misinterpretation errors frequently observed in existing DPO models. This design yields a 3.6% improvement over the strongest existing DPO baseline on visual question-answering tasks. To support future research, we release our complete framework, including all training data, model checkpoints, and our codebase at https://github.com/dmis-lab/med-vlm-dpo.
title Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2601.17918