RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Hanyang, Chen, Haoxian, Guo, Yucheng, Winata, Genta Indra, Ou, Tingting, Huang, Ziyu, Yao, David D., Tang, Wenpin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024)
by: Chen, Haoxian, et al.
Published: (2024)
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
by: Zhao, Hanyang, et al.
Published: (2024)
by: Zhao, Hanyang, et al.
Published: (2024)
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
Score-based Diffusion Models via Stochastic Differential Equations -- a Technical Tutorial
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
by: Sheng, Jiayuan, et al.
Published: (2025)
by: Sheng, Jiayuan, et al.
Published: (2025)
Scores as Actions: a framework of fine-tuning diffusion models by continuous-time reinforcement learning
by: Zhao, Hanyang, et al.
Published: (2024)
by: Zhao, Hanyang, et al.
Published: (2024)
MINERS: Multilingual Language Models as Semantic Retrievers
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Contractive Diffusion Probabilistic Models
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025)
by: Khelli, Maria, et al.
Published: (2025)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
by: Yi, Hongzhu, et al.
Published: (2026)
by: Yi, Hongzhu, et al.
Published: (2026)
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
by: Merin, Adril Putra, et al.
Published: (2026)
by: Merin, Adril Putra, et al.
Published: (2026)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
SOCRATES: Simulation Optimization with Correlated Replicas and Adaptive Trajectory Evaluations
by: Zhang, Haoting, et al.
Published: (2025)
by: Zhang, Haoting, et al.
Published: (2025)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation
by: Yan, Shi-Qi, et al.
Published: (2025)
by: Yan, Shi-Qi, et al.
Published: (2025)
Fine-tuning of diffusion models via stochastic control: entropy regularization and beyond
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
Do Language Models Understand Honorific Systems in Javanese?
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization
by: Gu, Yi, et al.
Published: (2024)
by: Gu, Yi, et al.
Published: (2024)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Retrieval-Augmented Fine-Tuning With Preference Optimization For Visual Program Generation
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
Diffusion Generative Models Meet Compressed Sensing, with Applications to Imaging and Finance
by: Guo, Zhengyi, et al.
Published: (2025)
by: Guo, Zhengyi, et al.
Published: (2025)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
by: Niu, Tianyi, et al.
Published: (2026)
by: Niu, Tianyi, et al.
Published: (2026)
Similar Items
-
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024) -
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026) -
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024) -
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
by: Zhao, Hanyang, et al.
Published: (2024) -
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)