Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Wenqi, Song, Xuemeng, Li, Jiaxi, Wei, Yinwei, Zheng, Na, Yin, Jianhua, Nie, Liqiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914213316788224
author Liu, Wenqi
Song, Xuemeng
Li, Jiaxi
Wei, Yinwei
Zheng, Na
Yin, Jianhua
Nie, Liqiang
author_facet Liu, Wenqi
Song, Xuemeng
Li, Jiaxi
Wei, Yinwei
Zheng, Na
Yin, Jianhua
Nie, Liqiang
contents Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visual inputs and hence reducing hallucination, they suffer from non-rigorous optimization objective function and indirect preference supervision. To address these limitations, we propose a Symmetric Multimodal Preference Optimization (SymMPO), which conducts symmetric preference learning with direct preference supervision (i.e., response pairs) for visual understanding enhancement, while maintaining rigorous theoretical alignment with standard DPO. In addition to conventional ordinal preference learning, SymMPO introduces a preference margin consistency loss to quantitatively regulate the preference gap between symmetric preference pairs. Comprehensive evaluation across five benchmarks demonstrate SymMPO's superior performance, validating its effectiveness in hallucination mitigation of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
Liu, Wenqi
Song, Xuemeng
Li, Jiaxi
Wei, Yinwei
Zheng, Na
Yin, Jianhua
Nie, Liqiang
Artificial Intelligence
Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visual inputs and hence reducing hallucination, they suffer from non-rigorous optimization objective function and indirect preference supervision. To address these limitations, we propose a Symmetric Multimodal Preference Optimization (SymMPO), which conducts symmetric preference learning with direct preference supervision (i.e., response pairs) for visual understanding enhancement, while maintaining rigorous theoretical alignment with standard DPO. In addition to conventional ordinal preference learning, SymMPO introduces a preference margin consistency loss to quantitatively regulate the preference gap between symmetric preference pairs. Comprehensive evaluation across five benchmarks demonstrate SymMPO's superior performance, validating its effectiveness in hallucination mitigation of MLLMs.
title Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
topic Artificial Intelligence
url https://arxiv.org/abs/2506.11712