Visual Preference Optimization with Rubric Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Ya-Qi, Hong, Fangyu, Qu, Xiangyang, Wang, Hao, Wu, Gaojie, Luo, Qiaoyu, Xu, Nuo, Wang, Huixin, Xu, Wuheng, Liao, Yongxin, Chen, Zihao, Li, Haonan, Li, Ziming, Peng, Dezhi, Liao, Minghui, Wu, Jihao, Ren, Haoyu, Tu, Dandan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918446426488832
author Yu, Ya-Qi
Hong, Fangyu
Qu, Xiangyang
Wang, Hao
Wu, Gaojie
Luo, Qiaoyu
Xu, Nuo
Wang, Huixin
Xu, Wuheng
Liao, Yongxin
Chen, Zihao
Li, Haonan
Li, Ziming
Peng, Dezhi
Liao, Minghui
Wu, Jihao
Ren, Haoyu
Tu, Dandan
author_facet Yu, Ya-Qi
Hong, Fangyu
Qu, Xiangyang
Wang, Hao
Wu, Gaojie
Luo, Qiaoyu
Xu, Nuo
Wang, Huixin
Xu, Wuheng
Liao, Yongxin
Chen, Zihao
Li, Haonan
Li, Ziming
Peng, Dezhi
Liao, Minghui
Wu, Jihao
Ren, Haoyu
Tu, Dandan
contents The effectiveness of Direct Preference Optimization (DPO) depends on preference data that reflect the quality differences that matter in multimodal tasks. Existing pipelines often rely on off-policy perturbations or coarse outcome-based signals, which are not well suited to fine-grained visual reasoning. We propose rDPO, a preference optimization framework based on instance-specific rubrics. For each image-instruction pair, we create a checklist-style rubric of essential and additional criteria to score responses from any possible policies. The instruction-rubric pool is built offline and reused during the construction of on-policy data. On public reward modeling benchmarks, rubric-based prompting massively improves a 30B-A3B judge and brings it close to GPT-5.4. On public downstream benchmarks, rubric-based filtering raises the macro average to 82.69, whereas outcome-based filtering drops it to 75.82 from 81.14. When evaluating scalability on a comprehensive benchmark, rDPO achieves 61.01, markedly outperforming the style-constrained baseline (52.36) and surpassing the 59.48 base model. Together, these results show that visual preference optimization benefits from combining on-policy data construction with instance-specific criterion-level feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13029
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Visual Preference Optimization with Rubric Rewards
Yu, Ya-Qi
Hong, Fangyu
Qu, Xiangyang
Wang, Hao
Wu, Gaojie
Luo, Qiaoyu
Xu, Nuo
Wang, Huixin
Xu, Wuheng
Liao, Yongxin
Chen, Zihao
Li, Haonan
Li, Ziming
Peng, Dezhi
Liao, Minghui
Wu, Jihao
Ren, Haoyu
Tu, Dandan
Computer Vision and Pattern Recognition
Artificial Intelligence
The effectiveness of Direct Preference Optimization (DPO) depends on preference data that reflect the quality differences that matter in multimodal tasks. Existing pipelines often rely on off-policy perturbations or coarse outcome-based signals, which are not well suited to fine-grained visual reasoning. We propose rDPO, a preference optimization framework based on instance-specific rubrics. For each image-instruction pair, we create a checklist-style rubric of essential and additional criteria to score responses from any possible policies. The instruction-rubric pool is built offline and reused during the construction of on-policy data. On public reward modeling benchmarks, rubric-based prompting massively improves a 30B-A3B judge and brings it close to GPT-5.4. On public downstream benchmarks, rubric-based filtering raises the macro average to 82.69, whereas outcome-based filtering drops it to 75.82 from 81.14. When evaluating scalability on a comprehensive benchmark, rDPO achieves 61.01, markedly outperforming the style-constrained baseline (52.36) and surpassing the 59.48 base model. Together, these results show that visual preference optimization benefits from combining on-policy data construction with instance-specific criterion-level feedback.
title Visual Preference Optimization with Rubric Rewards
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.13029