Continual SFT Matches Multimodal RLHF with Negative Supervision

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Ke, Wang, Yu, Sun, Yanpeng, Chen, Qiang, Liu, Jiangjiang, Zhang, Gang, Wang, Jingdong
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917844959100928
author Zhu, Ke
Wang, Yu
Sun, Yanpeng
Chen, Qiang
Liu, Jiangjiang
Zhang, Gang
Wang, Jingdong
author_facet Zhu, Ke
Wang, Yu
Sun, Yanpeng
Chen, Qiang
Liu, Jiangjiang
Zhang, Gang
Wang, Jingdong
contents Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this preference alignment stage. In this paper, we observe that the inherent value of multimodal RLHF lies in its negative supervision, the logit of the rejected responses. We thus propose a novel negative supervised finetuning (nSFT) approach that fully excavates these information resided. Our nSFT disentangles this negative supervision in RLHF paradigm, and continually aligns VLMs with a simple SFT loss. This is more memory efficient than multimodal RLHF where 2 (e.g., DPO) or 4 (e.g., PPO) large VLMs are strictly required. The effectiveness of nSFT is rigorously proved by comparing it with various multimodal RLHF approaches, across different dataset sources, base VLMs and evaluation metrics. Besides, fruitful of ablations are provided to support our hypothesis. We hope this paper will stimulate further research to properly align large vision language models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14797
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continual SFT Matches Multimodal RLHF with Negative Supervision
Zhu, Ke
Wang, Yu
Sun, Yanpeng
Chen, Qiang
Liu, Jiangjiang
Zhang, Gang
Wang, Jingdong
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this preference alignment stage. In this paper, we observe that the inherent value of multimodal RLHF lies in its negative supervision, the logit of the rejected responses. We thus propose a novel negative supervised finetuning (nSFT) approach that fully excavates these information resided. Our nSFT disentangles this negative supervision in RLHF paradigm, and continually aligns VLMs with a simple SFT loss. This is more memory efficient than multimodal RLHF where 2 (e.g., DPO) or 4 (e.g., PPO) large VLMs are strictly required. The effectiveness of nSFT is rigorously proved by comparing it with various multimodal RLHF approaches, across different dataset sources, base VLMs and evaluation metrics. Besides, fruitful of ablations are provided to support our hypothesis. We hope this paper will stimulate further research to properly align large vision language models.
title Continual SFT Matches Multimodal RLHF with Negative Supervision
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.14797