Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Diwan, Nirav, Ergen, Tolga, Shim, Dongsub, Lee, Honglak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
MapExplorer: New Content Generation from Low-Dimensional Visualizations
by: Zhang, Xingjian, et al.
Published: (2024)
by: Zhang, Xingjian, et al.
Published: (2024)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Preference Guided Iterated Pareto Referent Optimisation for Accessible Route Planning
by: Speziali, Paolo, et al.
Published: (2026)
by: Speziali, Paolo, et al.
Published: (2026)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
by: Zhang, Xingjian, et al.
Published: (2024)
by: Zhang, Xingjian, et al.
Published: (2024)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
by: Zhang, Lechen, et al.
Published: (2024)
by: Zhang, Lechen, et al.
Published: (2024)
Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
by: Wu, Skyler, et al.
Published: (2025)
by: Wu, Skyler, et al.
Published: (2025)
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2024)
by: Gao, Chen-Xiao, et al.
Published: (2024)
When Can Proxies Improve the Sample Complexity of Preference Learning?
by: Zhu, Yuchen, et al.
Published: (2024)
by: Zhu, Yuchen, et al.
Published: (2024)
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)
by: Yuan, Suqin, et al.
Published: (2026)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
by: Belakaria, Syrine, et al.
Published: (2025)
by: Belakaria, Syrine, et al.
Published: (2025)
Towards Diverse Evaluation of Class Incremental Learning: A Representation Learning Perspective
by: Cha, Sungmin, et al.
Published: (2022)
by: Cha, Sungmin, et al.
Published: (2022)
Debiasing Online Preference Learning via Preference Feature Preservation
by: Kim, Dongyoung, et al.
Published: (2025)
by: Kim, Dongyoung, et al.
Published: (2025)
LEASE: Offline Preference-based Reinforcement Learning with High Sample Efficiency
by: Liu, Xiao-Yin, et al.
Published: (2024)
by: Liu, Xiao-Yin, et al.
Published: (2024)
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
by: Metcalf, Katherine, et al.
Published: (2024)
by: Metcalf, Katherine, et al.
Published: (2024)
POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
Preference-Conditioned Language-Guided Abstraction
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
Pareto-Optimal Learning from Preferences with Hidden Context
by: Bahlous-Boldi, Ryan, et al.
Published: (2024)
by: Bahlous-Boldi, Ryan, et al.
Published: (2024)
TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
by: Chen, Ziyuan, et al.
Published: (2025)
by: Chen, Ziyuan, et al.
Published: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
by: Shianifar, Jonaid, et al.
Published: (2026)
by: Shianifar, Jonaid, et al.
Published: (2026)
Preference-Guided Diffusion for Multi-Objective Offline Optimization
by: Annadani, Yashas, et al.
Published: (2025)
by: Annadani, Yashas, et al.
Published: (2025)
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
by: Cheng, Jie, et al.
Published: (2024)
by: Cheng, Jie, et al.
Published: (2024)
GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
by: Zhao, Yiyang, et al.
Published: (2025)
by: Zhao, Yiyang, et al.
Published: (2025)
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
by: Mahmud, Saaduddin, et al.
Published: (2024)
by: Mahmud, Saaduddin, et al.
Published: (2024)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Sample Efficient Preference Alignment in LLMs via Active Exploration
by: Mehta, Viraj, et al.
Published: (2023)
by: Mehta, Viraj, et al.
Published: (2023)
Preference Elicitation for Offline Reinforcement Learning
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Reward Learning From Preference With Ties
by: Liu, Jinsong, et al.
Published: (2024)
by: Liu, Jinsong, et al.
Published: (2024)
Meta-Learning Objectives for Preference Optimization
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
Learning to Select In-Context Demonstration Preferred by Large Language Model
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning
by: Hiroshima, Kei, et al.
Published: (2026)
by: Hiroshima, Kei, et al.
Published: (2026)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Geometric-Averaged Preference Optimization for Soft Preference Labels
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences
by: Kim, Euihyun, et al.
Published: (2025)
by: Kim, Euihyun, et al.
Published: (2025)
Similar Items
-
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025) -
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025) -
MapExplorer: New Content Generation from Low-Dimensional Visualizations
by: Zhang, Xingjian, et al.
Published: (2024) -
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025) -
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)