AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Taneesh, Madhavan, Rahul, Zhang, Xuchao, Bansal, Chetan, Rajmohan, Saravan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
REFA: Reference Free Alignment for multi-preference optimization
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Anyprefer: An Agentic Framework for Preference Data Synthesis
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
SynthAgent: Adapting Web Agents with Synthetic Supervision
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
TSO: Self-Training with Scaled Preference Optimization
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
On the Role of Preference Variance in Preference Optimization
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Exploring LLM-based Agents for Root Cause Analysis
von: Roy, Devjeet, et al.
Veröffentlicht: (2024)
von: Roy, Devjeet, et al.
Veröffentlicht: (2024)
Active Preference Learning for Large Language Models
von: Muldrew, William, et al.
Veröffentlicht: (2024)
von: Muldrew, William, et al.
Veröffentlicht: (2024)
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
Geometric-Averaged Preference Optimization for Soft Preference Labels
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
Selective Preference Optimization via Token-Level Reward Function Estimation
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
von: Jian, Chengtao, et al.
Veröffentlicht: (2025)
von: Jian, Chengtao, et al.
Veröffentlicht: (2025)
CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2024)
Filtered Direct Preference Optimization
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2024)
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2024)
Direct Preference Optimization with an Offset
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Preference Learning Algorithms Do Not Learn Preference Rankings
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
Generative Caching for Structurally Similar Prompts and Responses
von: Chakraborty, Sarthak, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sarthak, et al.
Veröffentlicht: (2025)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
von: Melikidze, Davit, et al.
Veröffentlicht: (2026)
von: Melikidze, Davit, et al.
Veröffentlicht: (2026)
Entropy Controllable Direct Preference Optimization
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
Orthogonal Finetuning for Direct Preference Optimization
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
T-REG: Preference Optimization with Token-Level Reward Regularization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Active Preference Inference using Language Models and Probabilistic Reasoning
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
Preference Optimization by Estimating the Ratio of the Data Distribution
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
VPO: Leveraging the Number of Votes in Preference Optimization
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
Synergistic Weak-Strong Collaboration by Aligning Preferences
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024) -
REFA: Reference Free Alignment for multi-preference optimization
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024) -
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024) -
Anyprefer: An Agentic Framework for Preference Data Synthesis
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025) -
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)