Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Amirloo, Elmira, Fauconnier, Jean-Philippe, Roesmann, Christoph, Kerl, Christian, Boney, Rinu, Qian, Yusu, Wang, Zirui, Dehghan, Afshin, Yang, Yinfei, Gan, Zhe, Grasch, Peter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917710648049664
author Amirloo, Elmira
Fauconnier, Jean-Philippe
Roesmann, Christoph
Kerl, Christian
Boney, Rinu
Qian, Yusu
Wang, Zirui
Dehghan, Afshin
Yang, Yinfei
Gan, Zhe
Grasch, Peter
author_facet Amirloo, Elmira
Fauconnier, Jean-Philippe
Roesmann, Christoph
Kerl, Christian
Boney, Rinu
Qian, Yusu
Wang, Zirui
Dehghan, Afshin
Yang, Yinfei
Gan, Zhe
Grasch, Peter
contents Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to language models, MLLMs for image understanding tasks encounter challenges like hallucination. In MLLMs, hallucination can occur not only by stating incorrect facts but also by producing responses that are inconsistent with the image content. A primary objective of alignment for MLLMs is to encourage these models to align responses more closely with image information. Recently, multiple works have introduced preference datasets for MLLMs and examined different alignment methods, including Direct Preference Optimization (DPO) and Proximal Policy Optimization (PPO). However, due to variations in datasets, base model types, and alignment methods, it remains unclear which specific elements contribute most significantly to the reported improvements in these works. In this paper, we independently analyze each aspect of preference alignment in MLLMs. We start by categorizing the alignment algorithms into two groups, offline (such as DPO), and online (such as online-DPO), and show that combining offline and online methods can improve the performance of the model in certain scenarios. We review a variety of published multimodal preference datasets and discuss how the details of their construction impact model performance. Based on these insights, we introduce a novel way of creating multimodal preference data called Bias-Driven Hallucination Sampling (BDHS) that needs neither additional annotation nor external models, and show that it can achieve competitive performance to previously published alignment work for multimodal models across a range of benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02477
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Amirloo, Elmira
Fauconnier, Jean-Philippe
Roesmann, Christoph
Kerl, Christian
Boney, Rinu
Qian, Yusu
Wang, Zirui
Dehghan, Afshin
Yang, Yinfei
Gan, Zhe
Grasch, Peter
Computer Vision and Pattern Recognition
Computation and Language
Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to language models, MLLMs for image understanding tasks encounter challenges like hallucination. In MLLMs, hallucination can occur not only by stating incorrect facts but also by producing responses that are inconsistent with the image content. A primary objective of alignment for MLLMs is to encourage these models to align responses more closely with image information. Recently, multiple works have introduced preference datasets for MLLMs and examined different alignment methods, including Direct Preference Optimization (DPO) and Proximal Policy Optimization (PPO). However, due to variations in datasets, base model types, and alignment methods, it remains unclear which specific elements contribute most significantly to the reported improvements in these works. In this paper, we independently analyze each aspect of preference alignment in MLLMs. We start by categorizing the alignment algorithms into two groups, offline (such as DPO), and online (such as online-DPO), and show that combining offline and online methods can improve the performance of the model in certain scenarios. We review a variety of published multimodal preference datasets and discuss how the details of their construction impact model performance. Based on these insights, we introduce a novel way of creating multimodal preference data called Bias-Driven Hallucination Sampling (BDHS) that needs neither additional annotation nor external models, and show that it can achieve competitive performance to previously published alignment work for multimodal models across a range of benchmarks.
title Understanding Alignment in Multimodal LLMs: A Comprehensive Study
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2407.02477