Saved in:
Bibliographic Details
Main Authors: Lu, Yiting, Guan, Fengbin, Gao, Yixin, Zhong, Yan, Peng, Xinge, Yuan, Jiakang, Liu, Yihao, Zhang, Bo, Li, Xin, Chen, Zhibo, Lin, Weisi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.10609
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912644138532864
author Lu, Yiting
Guan, Fengbin
Gao, Yixin
Zhong, Yan
Peng, Xinge
Yuan, Jiakang
Liu, Yihao
Zhang, Bo
Li, Xin
Chen, Zhibo
Lin, Weisi
author_facet Lu, Yiting
Guan, Fengbin
Gao, Yixin
Zhong, Yan
Peng, Xinge
Yuan, Jiakang
Liu, Yihao
Zhang, Bo
Li, Xin
Chen, Zhibo
Lin, Weisi
contents Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable reward signals for policy optimization. Inspired by subjective experiments, where participants are given task-specific instructions outlining distinct assessment principles prior to evaluation, we propose OmniQuality-R, a structured reward modeling framework that transforms multi-dimensional reasoning into continuous and interpretable reward signals. To enable this, we construct a reasoning-enhanced reward modeling dataset by sampling informative plan-reason trajectories via rejection sampling, forming a reliable chain-of-thought (CoT) dataset for supervised fine-tuning (SFT). Building on this, we apply Group Relative Policy Optimization (GRPO) for post-training, using a Gaussian-based reward to support continuous score prediction. To further stabilize the training and improve downstream generalization, we incorporate standard deviation (STD) filtering and entropy gating mechanisms during reinforcement learning. These techniques suppress unstable updates and reduce variance in policy optimization. We evaluate OmniQuality-R on three key IQA tasks: aesthetic quality assessment, technical quality evaluation, and text-image alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10609
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
Lu, Yiting
Guan, Fengbin
Gao, Yixin
Zhong, Yan
Peng, Xinge
Yuan, Jiakang
Liu, Yihao
Zhang, Bo
Li, Xin
Chen, Zhibo
Lin, Weisi
Computer Vision and Pattern Recognition
Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable reward signals for policy optimization. Inspired by subjective experiments, where participants are given task-specific instructions outlining distinct assessment principles prior to evaluation, we propose OmniQuality-R, a structured reward modeling framework that transforms multi-dimensional reasoning into continuous and interpretable reward signals. To enable this, we construct a reasoning-enhanced reward modeling dataset by sampling informative plan-reason trajectories via rejection sampling, forming a reliable chain-of-thought (CoT) dataset for supervised fine-tuning (SFT). Building on this, we apply Group Relative Policy Optimization (GRPO) for post-training, using a Gaussian-based reward to support continuous score prediction. To further stabilize the training and improve downstream generalization, we incorporate standard deviation (STD) filtering and entropy gating mechanisms during reinforcement learning. These techniques suppress unstable updates and reduce variance in policy optimization. We evaluate OmniQuality-R on three key IQA tasks: aesthetic quality assessment, technical quality evaluation, and text-image alignment.
title OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10609