Listener-Rewarded Thinking in VLMs for Image Preferences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gambashidze, Alexander, Pengyi, Li, Skripkin, Matvey, Galichin, Andrey, Gusarov, Anton, Sobolev, Konstantin, Kuznetsov, Andrey, Oseledets, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917396918304768
author Gambashidze, Alexander
Pengyi, Li
Skripkin, Matvey
Galichin, Andrey
Gusarov, Anton
Sobolev, Konstantin
Kuznetsov, Andrey
Oseledets, Ivan
author_facet Gambashidze, Alexander
Pengyi, Li
Skripkin, Matvey
Galichin, Andrey
Gusarov, Anton
Sobolev, Konstantin
Kuznetsov, Andrey
Oseledets, Ivan
contents Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However, current reward models often fail to generalize, and supervised fine-tuning leads to memorization, demanding complex annotation pipelines. While reinforcement learning (RL), specifically Group Relative Policy Optimization (GRPO), improves generalization, we uncover a key failure mode: a significant drop in reasoning accuracy occurs when a model's reasoning trace contradicts that of an independent, frozen vision-language model ("listener") evaluating the same output. To address this, we introduce a listener-augmented GRPO framework. Here, the listener re-evaluates the reasoner's chain-of-thought to provide a dense, calibrated confidence score, shaping the RL reward signal. This encourages the reasoner not only to answer correctly, but to produce explanations that are persuasive to an independent model. Our listener-shaped reward scheme achieves best accuracy on the ImageReward benchmark (67.4%), significantly improves out-of-distribution (OOD) performance on a large-scale human preference dataset (1.2M votes, up to +6% over naive reasoner), and reduces reasoning contradictions compared to strong GRPO and SFT baselines. These results demonstrate that listener-based rewards provide a scalable, data-efficient path to aligning vision-language models with nuanced human preferences. We will release our reasoning model here: https://huggingface.co/alexgambashidze/qwen2.5vl_image_preference_reasoner.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listener-Rewarded Thinking in VLMs for Image Preferences
Gambashidze, Alexander
Pengyi, Li
Skripkin, Matvey
Galichin, Andrey
Gusarov, Anton
Sobolev, Konstantin
Kuznetsov, Andrey
Oseledets, Ivan
Computer Vision and Pattern Recognition
Artificial Intelligence
Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However, current reward models often fail to generalize, and supervised fine-tuning leads to memorization, demanding complex annotation pipelines. While reinforcement learning (RL), specifically Group Relative Policy Optimization (GRPO), improves generalization, we uncover a key failure mode: a significant drop in reasoning accuracy occurs when a model's reasoning trace contradicts that of an independent, frozen vision-language model ("listener") evaluating the same output. To address this, we introduce a listener-augmented GRPO framework. Here, the listener re-evaluates the reasoner's chain-of-thought to provide a dense, calibrated confidence score, shaping the RL reward signal. This encourages the reasoner not only to answer correctly, but to produce explanations that are persuasive to an independent model. Our listener-shaped reward scheme achieves best accuracy on the ImageReward benchmark (67.4%), significantly improves out-of-distribution (OOD) performance on a large-scale human preference dataset (1.2M votes, up to +6% over naive reasoner), and reduces reasoning contradictions compared to strong GRPO and SFT baselines. These results demonstrate that listener-based rewards provide a scalable, data-efficient path to aligning vision-language models with nuanced human preferences. We will release our reasoning model here: https://huggingface.co/alexgambashidze/qwen2.5vl_image_preference_reasoner.
title Listener-Rewarded Thinking in VLMs for Image Preferences
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.22832