OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Junzhe, Zhang, Tianshu, Huang, Shiyu, Niu, Yuwei, Sun, Chao, Zhang, Rongzhou, Zhou, Guanyu, Wen, Lijie, Hu, Xuming
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911131844476928
author Chen, Junzhe
Zhang, Tianshu
Huang, Shiyu
Niu, Yuwei
Sun, Chao
Zhang, Rongzhou
Zhou, Guanyu
Wen, Lijie
Hu, Xuming
author_facet Chen, Junzhe
Zhang, Tianshu
Huang, Shiyu
Niu, Yuwei
Sun, Chao
Zhang, Rongzhou
Zhou, Guanyu
Wen, Lijie
Hu, Xuming
contents Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text modality tend to dominate, leading OLLMs to rely more heavily on textual cues while neglecting visual and audio information. In addition, fully multimodal scenarios introduce new challenges. Most existing models align visual or auditory modalities with text independently during training, while ignoring the intrinsic correlations between video and its corresponding audio. This oversight results in hallucinations when reasoning requires interpreting hidden audio cues embedded in video content. To address these challenges, we propose OmniDPO, a preference-alignment framework designed to mitigate hallucinations in OLLMs. Specifically, OmniDPO incorporates two strategies: (1) constructing text-preference sample pairs to enhance the model's understanding of audio-video interactions; and (2) constructing multimodal-preference sample pairs to strengthen the model's attention to visual and auditory information. By tackling both challenges, OmniDPO effectively improves multimodal grounding and reduces hallucination. Experiments conducted on two OLLMs demonstrate that OmniDPO not only effectively mitigates multimodal hallucinations but also significantly enhances the models' reasoning capabilities across modalities. All code and datasets will be released upon paper acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
Chen, Junzhe
Zhang, Tianshu
Huang, Shiyu
Niu, Yuwei
Sun, Chao
Zhang, Rongzhou
Zhou, Guanyu
Wen, Lijie
Hu, Xuming
Artificial Intelligence
Multimedia
Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text modality tend to dominate, leading OLLMs to rely more heavily on textual cues while neglecting visual and audio information. In addition, fully multimodal scenarios introduce new challenges. Most existing models align visual or auditory modalities with text independently during training, while ignoring the intrinsic correlations between video and its corresponding audio. This oversight results in hallucinations when reasoning requires interpreting hidden audio cues embedded in video content. To address these challenges, we propose OmniDPO, a preference-alignment framework designed to mitigate hallucinations in OLLMs. Specifically, OmniDPO incorporates two strategies: (1) constructing text-preference sample pairs to enhance the model's understanding of audio-video interactions; and (2) constructing multimodal-preference sample pairs to strengthen the model's attention to visual and auditory information. By tackling both challenges, OmniDPO effectively improves multimodal grounding and reduces hallucination. Experiments conducted on two OLLMs demonstrate that OmniDPO not only effectively mitigates multimodal hallucinations but also significantly enhances the models' reasoning capabilities across modalities. All code and datasets will be released upon paper acceptance.
title OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
topic Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2509.00723