Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thanh Thi, Wilson, Campbell, Dalins, Janis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918137530679296
author Nguyen, Thanh Thi
Wilson, Campbell
Dalins, Janis
author_facet Nguyen, Thanh Thi
Wilson, Campbell
Dalins, Janis
contents Large Vision-Language Models (LVLMs) or multimodal large language models represent a significant advancement in artificial intelligence, enabling systems to understand and generate content across both visual and textual modalities. While large-scale pretraining has driven substantial progress, fine-tuning these models for aligning with human values or engaging in specific tasks or behaviors remains a critical challenge. Deep Reinforcement Learning (DRL) and Direct Preference Optimization (DPO) offer promising frameworks for this aligning process. While DRL enables models to optimize actions using reward signals instead of relying solely on supervised preference data, DPO directly aligns the policy with preferences, eliminating the need for an explicit reward model. This overview explores paradigms for fine-tuning LVLMs, highlighting how DRL and DPO techniques can be used to align models with human preferences and values, improve task performance, and enable adaptive multimodal interaction. We categorize key approaches, examine sources of preference data, reward signals, and discuss open challenges such as scalability, sample efficiency, continual learning, generalization, and safety. The goal is to provide a clear understanding of how DRL and DPO contribute to the evolution of robust and human-aligned LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06759
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
Nguyen, Thanh Thi
Wilson, Campbell
Dalins, Janis
Machine Learning
Artificial Intelligence
Large Vision-Language Models (LVLMs) or multimodal large language models represent a significant advancement in artificial intelligence, enabling systems to understand and generate content across both visual and textual modalities. While large-scale pretraining has driven substantial progress, fine-tuning these models for aligning with human values or engaging in specific tasks or behaviors remains a critical challenge. Deep Reinforcement Learning (DRL) and Direct Preference Optimization (DPO) offer promising frameworks for this aligning process. While DRL enables models to optimize actions using reward signals instead of relying solely on supervised preference data, DPO directly aligns the policy with preferences, eliminating the need for an explicit reward model. This overview explores paradigms for fine-tuning LVLMs, highlighting how DRL and DPO techniques can be used to align models with human preferences and values, improve task performance, and enable adaptive multimodal interaction. We categorize key approaches, examine sources of preference data, reward signals, and discuss open challenges such as scalability, sample efficiency, continual learning, generalization, and safety. The goal is to provide a clear understanding of how DRL and DPO contribute to the evolution of robust and human-aligned LVLMs.
title Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.06759