RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Hanyang, Winata, Genta Indra, Das, Anirban, Zhang, Shi-Xiong, Yao, David D., Tang, Wenpin, Sahu, Sambit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024)
by: Chen, Haoxian, et al.
Published: (2024)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
by: Chakraborty, Amartya, et al.
Published: (2025)
by: Chakraborty, Amartya, et al.
Published: (2025)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
MINERS: Multilingual Language Models as Semantic Retrievers
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Score-based Diffusion Models via Stochastic Differential Equations -- a Technical Tutorial
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Contractive Diffusion Probabilistic Models
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
by: Niu, Tianyi, et al.
Published: (2026)
by: Niu, Tianyi, et al.
Published: (2026)
DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
by: Ramnath, Sahana, et al.
Published: (2026)
by: Ramnath, Sahana, et al.
Published: (2026)
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
by: Merin, Adril Putra, et al.
Published: (2026)
by: Merin, Adril Putra, et al.
Published: (2026)
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025)
by: Khelli, Maria, et al.
Published: (2025)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
by: Sheng, Jiayuan, et al.
Published: (2025)
by: Sheng, Jiayuan, et al.
Published: (2025)
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Scores as Actions: a framework of fine-tuning diffusion models by continuous-time reinforcement learning
by: Zhao, Hanyang, et al.
Published: (2024)
by: Zhao, Hanyang, et al.
Published: (2024)
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Polynomial Voting Rules
by: Tang, Wenpin, et al.
Published: (2022)
by: Tang, Wenpin, et al.
Published: (2022)
SOCRATES: Simulation Optimization with Correlated Replicas and Adaptive Trajectory Evaluations
by: Zhang, Haoting, et al.
Published: (2025)
by: Zhang, Haoting, et al.
Published: (2025)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Do Language Models Understand Honorific Systems in Javanese?
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
Analysis of the Order Flow Auction under Proposer-Builder Separation on Blockchain
by: Ma, Ruofei, et al.
Published: (2025)
by: Ma, Ruofei, et al.
Published: (2025)
Your Model Diversity, Not Method, Determines Reasoning Strategy
by: Choraria, Moulik, et al.
Published: (2026)
by: Choraria, Moulik, et al.
Published: (2026)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning
by: Bhatt, Dwait, et al.
Published: (2026)
by: Bhatt, Dwait, et al.
Published: (2026)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline
by: Guo, Zhengyi, et al.
Published: (2026)
by: Guo, Zhengyi, et al.
Published: (2026)
Similar Items
-
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024) -
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026) -
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
by: Zhao, Hanyang, et al.
Published: (2025) -
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
by: Zhao, Bo, et al.
Published: (2025) -
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024)