Preference Optimization with Multi-Sample Comparisons
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chaoqi, Zhao, Zhuokai, Zhu, Chen, Sankararaman, Karthik Abinav, Valko, Michal, Cao, Xuefei, Chen, Zhaorun, Khabsa, Madian, Chen, Yuxin, Ma, Hao, Wang, Sinong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
On the Equivalence of Graph Convolution and Mixup
by: Han, Xiaotian, et al.
Published: (2023)
by: Han, Xiaotian, et al.
Published: (2023)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
by: Yu, Zishun, et al.
Published: (2025)
by: Yu, Zishun, et al.
Published: (2025)
The Perfect Blend: Redefining RLHF with Mixture of Judges
by: Xu, Tengyu, et al.
Published: (2024)
by: Xu, Tengyu, et al.
Published: (2024)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
by: He, Yun, et al.
Published: (2024)
by: He, Yun, et al.
Published: (2024)
Contextual Bandits with Packing and Covering Constraints: A Modular Lagrangian Approach via Regression
by: Slivkins, Aleksandrs, et al.
Published: (2022)
by: Slivkins, Aleksandrs, et al.
Published: (2022)
Direct Acquisition Optimization for Low-Budget Active Learning
by: Zhao, Zhuokai, et al.
Published: (2024)
by: Zhao, Zhuokai, et al.
Published: (2024)
Reinforcement Learning from User Feedback
by: Han, Eric, et al.
Published: (2025)
by: Han, Eric, et al.
Published: (2025)
High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuning
by: Franzmeyer, Tim, et al.
Published: (2025)
by: Franzmeyer, Tim, et al.
Published: (2025)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
by: Fang, Yixiong, et al.
Published: (2024)
by: Fang, Yixiong, et al.
Published: (2024)
AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
by: Qin, Chengwei, et al.
Published: (2025)
by: Qin, Chengwei, et al.
Published: (2025)
Safe Reinforcement Learning via Hierarchical Adaptive Chance-Constraint Safeguards
by: Chen, Zhaorun, et al.
Published: (2023)
by: Chen, Zhaorun, et al.
Published: (2023)
Bradley-Terry Policy Optimization for Generative Preference Modeling
by: Feng, Shengyu, et al.
Published: (2025)
by: Feng, Shengyu, et al.
Published: (2025)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
RankCLIP: Ranking-Consistent Language-Image Pretraining
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
GRAPE: Generalizing Robot Policy via Preference Alignment
by: Zhang, Zijian, et al.
Published: (2024)
by: Zhang, Zijian, et al.
Published: (2024)
Bandits on graphs and structures
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Accelerating PDE Surrogates via RL-Guided Mesh Optimization
by: Meng, Yang, et al.
Published: (2026)
by: Meng, Yang, et al.
Published: (2026)
Generalized Parallel Scaling with Interdependent Generations
by: Dong, Harry, et al.
Published: (2025)
by: Dong, Harry, et al.
Published: (2025)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Self-Evolving Multi-Agent Systems via Decentralized Memory
by: Hao, Guangya, et al.
Published: (2026)
by: Hao, Guangya, et al.
Published: (2026)
Token-Level LLM Collaboration via FusionRoute
by: Xiong, Nuoya, et al.
Published: (2026)
by: Xiong, Nuoya, et al.
Published: (2026)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
by: Yang, Xianjun, et al.
Published: (2025)
by: Yang, Xianjun, et al.
Published: (2025)
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
Optimization, Isoperimetric Inequalities, and Sampling via Lyapunov Potentials
by: Chen, August Y., et al.
Published: (2024)
by: Chen, August Y., et al.
Published: (2024)
Understanding the Impact of Sampling Quality in Direct Preference Optimization
by: Kim, Kyung Rok, et al.
Published: (2025)
by: Kim, Kyung Rok, et al.
Published: (2025)
Feature importance analysis for patient management decisions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Online combinatorial optimization with stochastic decision sets and adversarial losses
by: Neu, Gergely, et al.
Published: (2026)
by: Neu, Gergely, et al.
Published: (2026)
Distance metric learning for conditional anomaly detection
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Learning from a single labeled face and a stream of unlabeled data
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Revealing graph bandits for maximizing local influence
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Extreme bandits
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Reinforced Preference Optimization for Recommendation
by: Tan, Junfei, et al.
Published: (2025)
by: Tan, Junfei, et al.
Published: (2025)
Sample Complexity Bounds for Stochastic Shortest Path with a Generative Model
by: Tarbouriech, Jean, et al.
Published: (2026)
by: Tarbouriech, Jean, et al.
Published: (2026)
Boosting LLM Reasoning via Spontaneous Self-Correction
by: Zhao, Xutong, et al.
Published: (2025)
by: Zhao, Xutong, et al.
Published: (2025)
Active Policy Improvement from Multiple Black-box Oracles
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
by: Wang, Lirui, et al.
Published: (2025)
by: Wang, Lirui, et al.
Published: (2025)
Similar Items
-
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025) -
On the Equivalence of Graph Convolution and Mixup
by: Han, Xiaotian, et al.
Published: (2023) -
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
by: Yu, Zishun, et al.
Published: (2025) -
The Perfect Blend: Redefining RLHF with Mixture of Judges
by: Xu, Tengyu, et al.
Published: (2024) -
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
by: He, Yun, et al.
Published: (2024)