RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Hanyang, Chen, Haoxian, Guo, Yucheng, Winata, Genta Indra, Ou, Tingting, Huang, Ziyu, Yao, David D., Tang, Wenpin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909986127347712
author Zhao, Hanyang
Chen, Haoxian
Guo, Yucheng
Winata, Genta Indra
Ou, Tingting
Huang, Ziyu
Yao, David D.
Tang, Wenpin
author_facet Zhao, Hanyang
Chen, Haoxian
Guo, Yucheng
Winata, Genta Indra
Ou, Tingting
Huang, Ziyu
Yao, David D.
Tang, Wenpin
contents Traditional preference tuning methods for LLMs/Visual Generative Models often rely solely on reward model labeling, which can be opaque, offer limited insights into the rationale behind preferences, and are prone to issues such as reward hacking or overfitting. We introduce Rich Preference Optimization (RPO), a novel pipeline that leverages rich feedback signals from Vision Language Models (VLMs) to improve the curation of preference pairs for fine-tuning visual generative models like text-to-image diffusion models. Our approach begins with prompting VLMs to generate detailed critiques of synthesized images, from which we further prompt VLMs to extract reliable and actionable image editing instructions. By implementing these instructions, we create refined images, resulting in synthetic, informative preference pairs that serve as enhanced tuning datasets. We demonstrate the effectiveness of our pipeline and the resulting datasets in fine-tuning state-of-the-art diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11720
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
Zhao, Hanyang
Chen, Haoxian
Guo, Yucheng
Winata, Genta Indra
Ou, Tingting
Huang, Ziyu
Yao, David D.
Tang, Wenpin
Machine Learning
Artificial Intelligence
Traditional preference tuning methods for LLMs/Visual Generative Models often rely solely on reward model labeling, which can be opaque, offer limited insights into the rationale behind preferences, and are prone to issues such as reward hacking or overfitting. We introduce Rich Preference Optimization (RPO), a novel pipeline that leverages rich feedback signals from Vision Language Models (VLMs) to improve the curation of preference pairs for fine-tuning visual generative models like text-to-image diffusion models. Our approach begins with prompting VLMs to generate detailed critiques of synthesized images, from which we further prompt VLMs to extract reliable and actionable image editing instructions. By implementing these instructions, we create refined images, resulting in synthetic, informative preference pairs that serve as enhanced tuning datasets. We demonstrate the effectiveness of our pipeline and the resulting datasets in fine-tuning state-of-the-art diffusion models.
title RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.11720