mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Fei, Zhou, Wenxuan, Huang, James Y., Xu, Nan, Zhang, Sheng, Poon, Hoifung, Chen, Muhao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909338069630976
author Wang, Fei
Zhou, Wenxuan
Huang, James Y.
Xu, Nan
Zhang, Sheng
Poon, Hoifung
Chen, Muhao
author_facet Wang, Fei
Zhou, Wenxuan
Huang, James Y.
Xu, Nan
Zhang, Sheng
Poon, Hoifung
Chen, Muhao
contents Direct preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment. Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement. Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition. To address this problem, we propose mDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference. Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood -- an intrinsic problem of relative preference optimization. Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that mDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Wang, Fei
Zhou, Wenxuan
Huang, James Y.
Xu, Nan
Zhang, Sheng
Poon, Hoifung
Chen, Muhao
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Direct preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment. Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement. Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition. To address this problem, we propose mDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference. Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood -- an intrinsic problem of relative preference optimization. Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that mDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination.
title mDPO: Conditional Preference Optimization for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.11839