Why DPO is a Misspecified Estimator and How to Fix It
Fuente:
arXiv
Saved in:
| Main Authors: | Gopalan, Aditya, Chowdhury, Sayak Ray, Banerjee, Debangshu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
Relational DNN Verification With Cross Executional Bound Refinement
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
A Bayesian Approach for Discovering Time- Delayed Differential Equation from Data
by: Chowdhury, Debangshu, et al.
Published: (2025)
by: Chowdhury, Debangshu, et al.
Published: (2025)
Revisiting Social Welfare in Bandits: UCB is (Nearly) All You Need
by: Sarkar, Dhruv, et al.
Published: (2025)
by: Sarkar, Dhruv, et al.
Published: (2025)
DP-NCB: Privacy Preserving Fair Bandits
by: Sarkar, Dhruv, et al.
Published: (2025)
by: Sarkar, Dhruv, et al.
Published: (2025)
Improved Algorithms for Nash Welfare in Linear Bandits
by: Sarkar, Dhruv, et al.
Published: (2026)
by: Sarkar, Dhruv, et al.
Published: (2026)
Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
by: Bhattacharyya, Riddhiman, et al.
Published: (2026)
by: Bhattacharyya, Riddhiman, et al.
Published: (2026)
Data Shifts Hurt CoT: A Theoretical Study
by: Yin, Lang, et al.
Published: (2025)
by: Yin, Lang, et al.
Published: (2025)
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
by: Banerjee, Imon, et al.
Published: (2025)
by: Banerjee, Imon, et al.
Published: (2025)
Support is All You Need for Certified VAE Training
by: Xu, Changming, et al.
Published: (2025)
by: Xu, Changming, et al.
Published: (2025)
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
by: Son, Seongho, et al.
Published: (2024)
by: Son, Seongho, et al.
Published: (2024)
Why LLMs Cannot Think and How to Fix It
by: Jahrens, Marius, et al.
Published: (2025)
by: Jahrens, Marius, et al.
Published: (2025)
Active Preference Optimization for Sample Efficient RLHF
by: Das, Nirjhar, et al.
Published: (2024)
by: Das, Nirjhar, et al.
Published: (2024)
Preconditioned Robust Neural Posterior Estimation for Misspecified Simulators
by: Kelly, Ryan P., et al.
Published: (2026)
by: Kelly, Ryan P., et al.
Published: (2026)
Constrained Adversarial Perturbation
by: Nishad, Virendra, et al.
Published: (2025)
by: Nishad, Virendra, et al.
Published: (2025)
SETA: Statistical Fault Attribution for Compound AI Systems
by: Chowdhury, Sayak, et al.
Published: (2026)
by: Chowdhury, Sayak, et al.
Published: (2026)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
by: Pal, Arka, et al.
Published: (2024)
by: Pal, Arka, et al.
Published: (2024)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
On the Optimality of Misspecified Spectral Algorithms
by: Zhang, Haobo, et al.
Published: (2023)
by: Zhang, Haobo, et al.
Published: (2023)
CRANE: Reasoning with constrained LLM generation
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Inductive Domain Transfer In Misspecified Simulation-Based Inference
by: Senouf, Ortal, et al.
Published: (2025)
by: Senouf, Ortal, et al.
Published: (2025)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
by: Kaplan, Guy, et al.
Published: (2026)
by: Kaplan, Guy, et al.
Published: (2026)
DINGO: Constrained Inference for Diffusion LLMs
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
Incremental Randomized Smoothing Certification
by: Ugare, Shubham, et al.
Published: (2023)
by: Ugare, Shubham, et al.
Published: (2023)
Minimum mean-squared error estimation with bandit feedback
by: Ghosh, Ayon, et al.
Published: (2022)
by: Ghosh, Ayon, et al.
Published: (2022)
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)
by: Gopalan, Aditya, et al.
Published: (2022)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)
by: Maran, Davide, et al.
Published: (2026)
Spectral Algorithms in Misspecified Regression: Convergence under Covariate Shift
by: Liu, Ren-Rui, et al.
Published: (2025)
by: Liu, Ren-Rui, et al.
Published: (2025)
The Robustness of Differentiable Causal Discovery in Misspecified Scenarios
by: Yi, Huiyang, et al.
Published: (2025)
by: Yi, Huiyang, et al.
Published: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
How Global Calibration Strengthens Multiaccuracy
by: Casacuberta, Sílvia, et al.
Published: (2025)
by: Casacuberta, Sílvia, et al.
Published: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025)
by: Zhang, Zhengze, et al.
Published: (2025)
Information-Preserving Domain Transfer with Unlabeled Data in Misspecified Simulation-Based Inference
by: Jang, Joon, et al.
Published: (2026)
by: Jang, Joon, et al.
Published: (2026)
Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It
by: Xia, Guoxuan, et al.
Published: (2024)
by: Xia, Guoxuan, et al.
Published: (2024)
Similar Items
-
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023) -
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024) -
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025) -
Provably Robust DPO: Aligning Language Models with Noisy Feedback
by: Chowdhury, Sayak Ray, et al.
Published: (2024) -
Relational DNN Verification With Cross Executional Bound Refinement
by: Banerjee, Debangshu, et al.
Published: (2024)