VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Jipeng, Miao, Kehao, Pi, Renjie, Wang, Zhaowei, Liu, Runtao, Pan, Rui, Zhang, Tong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916796253077504
author Zhang, Jipeng
Miao, Kehao
Pi, Renjie
Wang, Zhaowei
Liu, Runtao
Pan, Rui
Zhang, Tong
author_facet Zhang, Jipeng
Miao, Kehao
Pi, Renjie
Wang, Zhaowei
Liu, Runtao
Pan, Rui
Zhang, Tong
contents Reinforcement Fine-Tuning (RFT) with verifiable rewards has advanced large language models but remains underexplored for Vision-Language (VL) models. The Vision-Language Reward Model (VL-RM) is key to aligning VL models by providing structured feedback, yet training effective VL-RMs faces two major challenges. First, the bootstrapping dilemma arises as high-quality training data depends on already strong VL models, creating a cycle where self-generated supervision reinforces existing biases. Second, modality bias and negative example amplification occur when VL models hallucinate incorrect visual attributes, leading to flawed preference data that further misguides training. To address these issues, we propose an iterative training framework leveraging vision experts, Chain-of-Thought (CoT) rationales, and Margin-based Rejection Sampling. Our approach refines preference datasets, enhances structured critiques, and iteratively improves reasoning. Experiments across VL-RM benchmarks demonstrate superior performance in hallucination detection and multimodal reasoning, advancing VL model alignment with reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13888
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
Zhang, Jipeng
Miao, Kehao
Pi, Renjie
Wang, Zhaowei
Liu, Runtao
Pan, Rui
Zhang, Tong
Computation and Language
Computer Vision and Pattern Recognition
Reinforcement Fine-Tuning (RFT) with verifiable rewards has advanced large language models but remains underexplored for Vision-Language (VL) models. The Vision-Language Reward Model (VL-RM) is key to aligning VL models by providing structured feedback, yet training effective VL-RMs faces two major challenges. First, the bootstrapping dilemma arises as high-quality training data depends on already strong VL models, creating a cycle where self-generated supervision reinforces existing biases. Second, modality bias and negative example amplification occur when VL models hallucinate incorrect visual attributes, leading to flawed preference data that further misguides training. To address these issues, we propose an iterative training framework leveraging vision experts, Chain-of-Thought (CoT) rationales, and Margin-based Rejection Sampling. Our approach refines preference datasets, enhances structured critiques, and iteratively improves reasoning. Experiments across VL-RM benchmarks demonstrate superior performance in hallucination detection and multimodal reasoning, advancing VL model alignment with reinforcement learning.
title VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13888