RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909986127347712 |
|---|---|
| author | Zhao, Hanyang Chen, Haoxian Guo, Yucheng Winata, Genta Indra Ou, Tingting Huang, Ziyu Yao, David D. Tang, Wenpin |
| author_facet | Zhao, Hanyang Chen, Haoxian Guo, Yucheng Winata, Genta Indra Ou, Tingting Huang, Ziyu Yao, David D. Tang, Wenpin |
| contents | Traditional preference tuning methods for LLMs/Visual Generative Models often rely solely on reward model labeling, which can be opaque, offer limited insights into the rationale behind preferences, and are prone to issues such as reward hacking or overfitting. We introduce Rich Preference Optimization (RPO), a novel pipeline that leverages rich feedback signals from Vision Language Models (VLMs) to improve the curation of preference pairs for fine-tuning visual generative models like text-to-image diffusion models. Our approach begins with prompting VLMs to generate detailed critiques of synthesized images, from which we further prompt VLMs to extract reliable and actionable image editing instructions. By implementing these instructions, we create refined images, resulting in synthetic, informative preference pairs that serve as enhanced tuning datasets. We demonstrate the effectiveness of our pipeline and the resulting datasets in fine-tuning state-of-the-art diffusion models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_11720 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences Zhao, Hanyang Chen, Haoxian Guo, Yucheng Winata, Genta Indra Ou, Tingting Huang, Ziyu Yao, David D. Tang, Wenpin Machine Learning Artificial Intelligence Traditional preference tuning methods for LLMs/Visual Generative Models often rely solely on reward model labeling, which can be opaque, offer limited insights into the rationale behind preferences, and are prone to issues such as reward hacking or overfitting. We introduce Rich Preference Optimization (RPO), a novel pipeline that leverages rich feedback signals from Vision Language Models (VLMs) to improve the curation of preference pairs for fine-tuning visual generative models like text-to-image diffusion models. Our approach begins with prompting VLMs to generate detailed critiques of synthesized images, from which we further prompt VLMs to extract reliable and actionable image editing instructions. By implementing these instructions, we create refined images, resulting in synthetic, informative preference pairs that serve as enhanced tuning datasets. We demonstrate the effectiveness of our pipeline and the resulting datasets in fine-tuning state-of-the-art diffusion models. |
| title | RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2503.11720 |