Video Generation Models Are Good Latent Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mi, Xiaoyue, Yu, Wenqing, Lian, Jiesong, Jie, Shibo, Zhong, Ruizhe, Liu, Zijun, Zhang, Guozhen, Zhou, Zixiang, Xu, Zhiyong, Zhou, Yuan, Lu, Qinglin, Tang, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917191241170944
author Mi, Xiaoyue
Yu, Wenqing
Lian, Jiesong
Jie, Shibo
Zhong, Ruizhe
Liu, Zijun
Zhang, Guozhen
Zhou, Zixiang
Xu, Zhiyong
Zhou, Yuan
Lu, Qinglin
Tang, Fan
author_facet Mi, Xiaoyue
Yu, Wenqing
Lian, Jiesong
Jie, Shibo
Zhong, Ruizhe
Liu, Zijun
Zhang, Guozhen
Zhou, Zixiang
Xu, Zhiyong
Zhou, Yuan
Lu, Qinglin
Tang, Fan
contents Reward feedback learning (ReFL) has proven effective for aligning image generation with human preferences. However, its extension to video generation faces significant challenges. Existing video reward models rely on vision-language models designed for pixel-space inputs, confining ReFL optimization to near-complete denoising steps after computationally expensive VAE decoding. This pixel-space approach incurs substantial memory overhead and increased training time, and its late-stage optimization lacks early-stage supervision, refining only visual quality rather than fundamental motion dynamics and structural coherence. In this work, we show that pre-trained video generation models are naturally suited for reward modeling in the noisy latent space, as they are explicitly designed to process noisy latent representations at arbitrary timesteps and inherently preserve temporal information through their sequential modeling capabilities. Accordingly, we propose Process Reward Feedback Learning~(PRFL), a framework that conducts preference optimization entirely in latent space, enabling efficient gradient backpropagation throughout the full denoising chain without VAE decoding. Extensive experiments demonstrate that PRFL significantly improves alignment with human preferences, while achieving substantial reductions in memory consumption and training time compared to RGB ReFL.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Generation Models Are Good Latent Reward Models
Mi, Xiaoyue
Yu, Wenqing
Lian, Jiesong
Jie, Shibo
Zhong, Ruizhe
Liu, Zijun
Zhang, Guozhen
Zhou, Zixiang
Xu, Zhiyong
Zhou, Yuan
Lu, Qinglin
Tang, Fan
Computer Vision and Pattern Recognition
Reward feedback learning (ReFL) has proven effective for aligning image generation with human preferences. However, its extension to video generation faces significant challenges. Existing video reward models rely on vision-language models designed for pixel-space inputs, confining ReFL optimization to near-complete denoising steps after computationally expensive VAE decoding. This pixel-space approach incurs substantial memory overhead and increased training time, and its late-stage optimization lacks early-stage supervision, refining only visual quality rather than fundamental motion dynamics and structural coherence. In this work, we show that pre-trained video generation models are naturally suited for reward modeling in the noisy latent space, as they are explicitly designed to process noisy latent representations at arbitrary timesteps and inherently preserve temporal information through their sequential modeling capabilities. Accordingly, we propose Process Reward Feedback Learning~(PRFL), a framework that conducts preference optimization entirely in latent space, enabling efficient gradient backpropagation throughout the full denoising chain without VAE decoding. Extensive experiments demonstrate that PRFL significantly improves alignment with human preferences, while achieving substantial reductions in memory consumption and training time compared to RGB ReFL.
title Video Generation Models Are Good Latent Reward Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21541