Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
Fuente:
arXiv
Saved in:
| Main Authors: | Im, Shawn, Li, Sharon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Well Can Preference Optimization Generalize Under Noisy Feedback?
by: Im, Shawn, et al.
Published: (2025)
by: Im, Shawn, et al.
Published: (2025)
A Unified Understanding and Evaluation of Steering Methods
by: Im, Shawn, et al.
Published: (2025)
by: Im, Shawn, et al.
Published: (2025)
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
by: Li, Wendi, et al.
Published: (2026)
by: Li, Wendi, et al.
Published: (2026)
Understanding the Learning Dynamics of Alignment with Human Feedback
by: Im, Shawn, et al.
Published: (2024)
by: Im, Shawn, et al.
Published: (2024)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
by: Im, Shawn, et al.
Published: (2026)
by: Im, Shawn, et al.
Published: (2026)
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
by: Oh, Changdae, et al.
Published: (2025)
by: Oh, Changdae, et al.
Published: (2025)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Scaling Laws for the Value of Individual Data Points in Machine Learning
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
Graph-SND: Sparse Aggregation for Behavioral Diversity in Multi-Agent Reinforcement Learning
by: Ray, Shawn
Published: (2026)
by: Ray, Shawn
Published: (2026)
Information-Theoretic Foundations for Neural Scaling Laws
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Theoretical Foundations of Scaling Law in Familial Models
by: Song, Huan, et al.
Published: (2025)
by: Song, Huan, et al.
Published: (2025)
Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
by: Zhou, Xingyu, et al.
Published: (2025)
by: Zhou, Xingyu, et al.
Published: (2025)
Can Language Models Discover Scaling Laws?
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
by: Yan, Tingkai, et al.
Published: (2025)
by: Yan, Tingkai, et al.
Published: (2025)
A Theoretical Framework for Explaining Reinforcement Learning with Shapley Values
by: Beechey, Daniel, et al.
Published: (2025)
by: Beechey, Daniel, et al.
Published: (2025)
Scaling Laws for Uncertainty in Deep Learning
by: Rosso, Mattia, et al.
Published: (2025)
by: Rosso, Mattia, et al.
Published: (2025)
Predictive AI Can Support Human Learning while Preserving Error Diversity
by: He, Vivianna Fang, et al.
Published: (2025)
by: He, Vivianna Fang, et al.
Published: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025)
by: Zhang, Zhengze, et al.
Published: (2025)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
by: Shi, Ruizhe, et al.
Published: (2025)
by: Shi, Ruizhe, et al.
Published: (2025)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Preference Robustness for DPO with Applications to Public Health
by: Kim, Cheol Woo, et al.
Published: (2025)
by: Kim, Cheol Woo, et al.
Published: (2025)
Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
by: Kamigaito, Hidetaka, et al.
Published: (2025)
by: Kamigaito, Hidetaka, et al.
Published: (2025)
Reranking Laws for Language Generation: A Communication-Theoretic Perspective
by: Farinhas, António, et al.
Published: (2024)
by: Farinhas, António, et al.
Published: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
Why DPO is a Misspecified Estimator and How to Fix It
by: Gopalan, Aditya, et al.
Published: (2025)
by: Gopalan, Aditya, et al.
Published: (2025)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
by: Pipano, Idan, et al.
Published: (2026)
by: Pipano, Idan, et al.
Published: (2026)
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
Reinforcement Learning from Diverse Human Preferences
by: Xue, Wanqi, et al.
Published: (2023)
by: Xue, Wanqi, et al.
Published: (2023)
Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
by: Li, Binghui, et al.
Published: (2025)
by: Li, Binghui, et al.
Published: (2025)
Scaling Laws are Redundancy Laws
by: Bi, Yuda, et al.
Published: (2025)
by: Bi, Yuda, et al.
Published: (2025)
LAD: Learning Advantage Distribution for Reasoning
by: Li, Wendi, et al.
Published: (2026)
by: Li, Wendi, et al.
Published: (2026)
Learning Human-like Representations to Enable Learning Human Values
by: Wynn, Andrea, et al.
Published: (2023)
by: Wynn, Andrea, et al.
Published: (2023)
Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced Data
by: Cortes, Corinna, et al.
Published: (2025)
by: Cortes, Corinna, et al.
Published: (2025)
Similar Items
-
How Well Can Preference Optimization Generalize Under Noisy Feedback?
by: Im, Shawn, et al.
Published: (2025) -
A Unified Understanding and Evaluation of Steering Methods
by: Im, Shawn, et al.
Published: (2025) -
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
by: Li, Wendi, et al.
Published: (2026) -
Understanding the Learning Dynamics of Alignment with Human Feedback
by: Im, Shawn, et al.
Published: (2024) -
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
by: Im, Shawn, et al.
Published: (2026)