Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Seongho, Bankes, William, Chowdhury, Sayak Ray, Paige, Brooks, Bogunovic, Ilija |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Multi-Objective Controlled Decoding of Large Language Models
by: Son, Seongho, et al.
Published: (2025)
by: Son, Seongho, et al.
Published: (2025)
REDUCR: Robust Data Downsampling Using Class Priority Reweighting
by: Bankes, William, et al.
Published: (2023)
by: Bankes, William, et al.
Published: (2023)
Active Preference Optimization for Sample Efficient RLHF
by: Das, Nirjhar, et al.
Published: (2024)
by: Das, Nirjhar, et al.
Published: (2024)
RSPO: Regularized Self-Play Alignment of Large Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
by: Ziomek, Juliusz, et al.
Published: (2026)
by: Ziomek, Juliusz, et al.
Published: (2026)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
Overton Pluralistic Reinforcement Learning for Large Language Models
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Sample-efficient Bayesian Optimisation Using Known Invariances
by: Brown, Theodore, et al.
Published: (2024)
by: Brown, Theodore, et al.
Published: (2024)
Robust Bayesian Optimisation with Unbounded Corruptions
by: Ezzerg, Abdelhamid, et al.
Published: (2025)
by: Ezzerg, Abdelhamid, et al.
Published: (2025)
Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2024)
by: Chidambaram, Keertana, et al.
Published: (2024)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
Sample Efficient Preference Alignment in LLMs via Active Exploration
by: Mehta, Viraj, et al.
Published: (2023)
by: Mehta, Viraj, et al.
Published: (2023)
Distributed Direct Preference Optimization
by: Jiang, Zhanhong
Published: (2026)
by: Jiang, Zhanhong
Published: (2026)
No-Regret Linear Bandits under Gap-Adjusted Misspecification
by: Liu, Chong, et al.
Published: (2025)
by: Liu, Chong, et al.
Published: (2025)
Gradient Imbalance in Direct Preference Optimization
by: Ma, Qinwei, et al.
Published: (2025)
by: Ma, Qinwei, et al.
Published: (2025)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
Lightweight Robust Direct Preference Optimization
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Why DPO is a Misspecified Estimator and How to Fix It
by: Gopalan, Aditya, et al.
Published: (2025)
by: Gopalan, Aditya, et al.
Published: (2025)
Improved Algorithms for Nash Welfare in Linear Bandits
by: Sarkar, Dhruv, et al.
Published: (2026)
by: Sarkar, Dhruv, et al.
Published: (2026)
Revisiting Social Welfare in Bandits: UCB is (Nearly) All You Need
by: Sarkar, Dhruv, et al.
Published: (2025)
by: Sarkar, Dhruv, et al.
Published: (2025)
DP-NCB: Privacy Preserving Fair Bandits
by: Sarkar, Dhruv, et al.
Published: (2025)
by: Sarkar, Dhruv, et al.
Published: (2025)
Drifting Preference Optimization for One-Step Generative Models
by: Jiang, Zhou, et al.
Published: (2026)
by: Jiang, Zhou, et al.
Published: (2026)
Resilient Contrastive Pre-training under Non-Stationary Drift
by: Yang, Xiaoyu, et al.
Published: (2025)
by: Yang, Xiaoyu, et al.
Published: (2025)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Risk-aware Direct Preference Optimization under Nested Risk Measure
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Continuous-Utility Direct Preference Optimization
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
Adversarially Robust Decision Transformer
by: Tang, Xiaohang, et al.
Published: (2024)
by: Tang, Xiaohang, et al.
Published: (2024)
Mean-Field Bayesian Optimisation
by: Steinberg, Petar, et al.
Published: (2025)
by: Steinberg, Petar, et al.
Published: (2025)
Uncertainty-Penalized Direct Preference Optimization
by: Houliston, Sam, et al.
Published: (2024)
by: Houliston, Sam, et al.
Published: (2024)
Direct Preference Optimization for Adaptive Concept-based Explanations
by: Teneggi, Jacopo, et al.
Published: (2025)
by: Teneggi, Jacopo, et al.
Published: (2025)
Understanding the Impact of Sampling Quality in Direct Preference Optimization
by: Kim, Kyung Rok, et al.
Published: (2025)
by: Kim, Kyung Rok, et al.
Published: (2025)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Accelerating Direct Preference Optimization with Prefix Sharing
by: Wang, Franklin, et al.
Published: (2024)
by: Wang, Franklin, et al.
Published: (2024)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Similar Items
-
Robust Multi-Objective Controlled Decoding of Large Language Models
by: Son, Seongho, et al.
Published: (2025) -
REDUCR: Robust Data Downsampling Using Class Priority Reweighting
by: Bankes, William, et al.
Published: (2023) -
Active Preference Optimization for Sample Efficient RLHF
by: Das, Nirjhar, et al.
Published: (2024) -
RSPO: Regularized Self-Play Alignment of Large Language Models
by: Tang, Xiaohang, et al.
Published: (2025) -
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
by: Ziomek, Juliusz, et al.
Published: (2026)