SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Jiesong, Zhong, Ruizhe, Zhou, Zixiang, Mi, Xiaoyue, Hu, Long, Zhou, Yuan, Lu, Qinglin, Hao, Yixue, Yan, Junchi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912966822068224
author Lian, Jiesong
Zhong, Ruizhe
Zhou, Zixiang
Mi, Xiaoyue
Hu, Long
Zhou, Yuan
Lu, Qinglin
Hao, Yixue
Yan, Junchi
author_facet Lian, Jiesong
Zhong, Ruizhe
Zhou, Zixiang
Mi, Xiaoyue
Hu, Long
Zhou, Yuan
Lu, Qinglin
Hao, Yixue
Yan, Junchi
contents Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodological hurdles. Current data collection paradigms, reliant on in-prompt pairwise annotations, suffer from labeling noise. Concurrently, the architectural design of VLM-based RMs, particularly their output mechanisms, remains underexplored. Furthermore, RM is susceptible to reward hacking in post-training. To mitigate these limitations, we propose SoliReward, a systematic framework for video RM training. Our framework first sources high-quality, cost-efficient data via single-item binary annotations, then constructs preference pairs using a cross-prompt pairing strategy. Architecturally, we employ a Hierarchical Progressive Query Attention mechanism to enhance feature aggregation. Finally, we introduce a modified BT loss that explicitly accommodates win-tie scenarios. This approach regularizes the RM's score distribution for positive samples, providing more nuanced preference signals to alleviate over-focus on a small number of top-scoring samples. Our approach is validated on benchmarks evaluating physical plausibility, subject deformity, and semantic alignment, demonstrating improvements in direct RM evaluation metrics and in the efficacy of post-training on video generation models. Code and benchmark are available at https://github.com/lian700/SoliReward.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22170
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
Lian, Jiesong
Zhong, Ruizhe
Zhou, Zixiang
Mi, Xiaoyue
Hu, Long
Zhou, Yuan
Lu, Qinglin
Hao, Yixue
Yan, Junchi
Machine Learning
Computer Vision and Pattern Recognition
Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodological hurdles. Current data collection paradigms, reliant on in-prompt pairwise annotations, suffer from labeling noise. Concurrently, the architectural design of VLM-based RMs, particularly their output mechanisms, remains underexplored. Furthermore, RM is susceptible to reward hacking in post-training. To mitigate these limitations, we propose SoliReward, a systematic framework for video RM training. Our framework first sources high-quality, cost-efficient data via single-item binary annotations, then constructs preference pairs using a cross-prompt pairing strategy. Architecturally, we employ a Hierarchical Progressive Query Attention mechanism to enhance feature aggregation. Finally, we introduce a modified BT loss that explicitly accommodates win-tie scenarios. This approach regularizes the RM's score distribution for positive samples, providing more nuanced preference signals to alleviate over-focus on a small number of top-scoring samples. Our approach is validated on benchmarks evaluating physical plausibility, subject deformity, and semantic alignment, demonstrating improvements in direct RM evaluation metrics and in the efficacy of post-training on video generation models. Code and benchmark are available at https://github.com/lian700/SoliReward.
title SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.22170