Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
Fuente:
arXiv
Salvato in:
| Autori principali: | Im, Shawn, Li, Sharon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Well Can Preference Optimization Generalize Under Noisy Feedback?
di: Im, Shawn, et al.
Pubblicazione: (2025)
di: Im, Shawn, et al.
Pubblicazione: (2025)
A Unified Understanding and Evaluation of Steering Methods
di: Im, Shawn, et al.
Pubblicazione: (2025)
di: Im, Shawn, et al.
Pubblicazione: (2025)
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
di: Li, Wendi, et al.
Pubblicazione: (2026)
di: Li, Wendi, et al.
Pubblicazione: (2026)
Understanding the Learning Dynamics of Alignment with Human Feedback
di: Im, Shawn, et al.
Pubblicazione: (2024)
di: Im, Shawn, et al.
Pubblicazione: (2024)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
di: Im, Shawn, et al.
Pubblicazione: (2026)
di: Im, Shawn, et al.
Pubblicazione: (2026)
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
di: Oh, Changdae, et al.
Pubblicazione: (2025)
di: Oh, Changdae, et al.
Pubblicazione: (2025)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
di: Mehta, Sushant, et al.
Pubblicazione: (2025)
di: Mehta, Sushant, et al.
Pubblicazione: (2025)
Scaling Laws for the Value of Individual Data Points in Machine Learning
di: Covert, Ian, et al.
Pubblicazione: (2024)
di: Covert, Ian, et al.
Pubblicazione: (2024)
Graph-SND: Sparse Aggregation for Behavioral Diversity in Multi-Agent Reinforcement Learning
di: Ray, Shawn
Pubblicazione: (2026)
di: Ray, Shawn
Pubblicazione: (2026)
Information-Theoretic Foundations for Neural Scaling Laws
di: Jeon, Hong Jun, et al.
Pubblicazione: (2024)
di: Jeon, Hong Jun, et al.
Pubblicazione: (2024)
Theoretical Foundations of Scaling Law in Familial Models
di: Song, Huan, et al.
Pubblicazione: (2025)
di: Song, Huan, et al.
Pubblicazione: (2025)
Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
di: Oldfield, James, et al.
Pubblicazione: (2025)
di: Oldfield, James, et al.
Pubblicazione: (2025)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
di: Su, Xuerui, et al.
Pubblicazione: (2025)
di: Su, Xuerui, et al.
Pubblicazione: (2025)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
di: Zhou, Xingyu, et al.
Pubblicazione: (2025)
di: Zhou, Xingyu, et al.
Pubblicazione: (2025)
Can Language Models Discover Scaling Laws?
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
How Feature Learning Can Improve Neural Scaling Laws
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
di: Li, Hongkang, et al.
Pubblicazione: (2025)
di: Li, Hongkang, et al.
Pubblicazione: (2025)
Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study
di: Liu, Hao, et al.
Pubblicazione: (2024)
di: Liu, Hao, et al.
Pubblicazione: (2024)
Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
di: Yan, Tingkai, et al.
Pubblicazione: (2025)
di: Yan, Tingkai, et al.
Pubblicazione: (2025)
A Theoretical Framework for Explaining Reinforcement Learning with Shapley Values
di: Beechey, Daniel, et al.
Pubblicazione: (2025)
di: Beechey, Daniel, et al.
Pubblicazione: (2025)
Scaling Laws for Uncertainty in Deep Learning
di: Rosso, Mattia, et al.
Pubblicazione: (2025)
di: Rosso, Mattia, et al.
Pubblicazione: (2025)
Predictive AI Can Support Human Learning while Preserving Error Diversity
di: He, Vivianna Fang, et al.
Pubblicazione: (2025)
di: He, Vivianna Fang, et al.
Pubblicazione: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
di: Shi, Ruizhe, et al.
Pubblicazione: (2025)
di: Shi, Ruizhe, et al.
Pubblicazione: (2025)
What Matters in Data for DPO?
di: Pan, Yu, et al.
Pubblicazione: (2025)
di: Pan, Yu, et al.
Pubblicazione: (2025)
Preference Robustness for DPO with Applications to Public Health
di: Kim, Cheol Woo, et al.
Pubblicazione: (2025)
di: Kim, Cheol Woo, et al.
Pubblicazione: (2025)
Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
di: Li, Binghui, et al.
Pubblicazione: (2024)
di: Li, Binghui, et al.
Pubblicazione: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
di: Kamigaito, Hidetaka, et al.
Pubblicazione: (2025)
di: Kamigaito, Hidetaka, et al.
Pubblicazione: (2025)
Reranking Laws for Language Generation: A Communication-Theoretic Perspective
di: Farinhas, António, et al.
Pubblicazione: (2024)
di: Farinhas, António, et al.
Pubblicazione: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
di: He, Chaoyue, et al.
Pubblicazione: (2026)
di: He, Chaoyue, et al.
Pubblicazione: (2026)
Why DPO is a Misspecified Estimator and How to Fix It
di: Gopalan, Aditya, et al.
Pubblicazione: (2025)
di: Gopalan, Aditya, et al.
Pubblicazione: (2025)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
di: Pipano, Idan, et al.
Pubblicazione: (2026)
di: Pipano, Idan, et al.
Pubblicazione: (2026)
It Takes Two: Your GRPO Is Secretly DPO
di: Wu, Yihong, et al.
Pubblicazione: (2025)
di: Wu, Yihong, et al.
Pubblicazione: (2025)
Robust Multi-Objective Preference Alignment with Online DPO
di: Gupta, Raghav, et al.
Pubblicazione: (2025)
di: Gupta, Raghav, et al.
Pubblicazione: (2025)
Reinforcement Learning from Diverse Human Preferences
di: Xue, Wanqi, et al.
Pubblicazione: (2023)
di: Xue, Wanqi, et al.
Pubblicazione: (2023)
Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
di: Li, Binghui, et al.
Pubblicazione: (2025)
di: Li, Binghui, et al.
Pubblicazione: (2025)
Scaling Laws are Redundancy Laws
di: Bi, Yuda, et al.
Pubblicazione: (2025)
di: Bi, Yuda, et al.
Pubblicazione: (2025)
LAD: Learning Advantage Distribution for Reasoning
di: Li, Wendi, et al.
Pubblicazione: (2026)
di: Li, Wendi, et al.
Pubblicazione: (2026)
Learning Human-like Representations to Enable Learning Human Values
di: Wynn, Andrea, et al.
Pubblicazione: (2023)
di: Wynn, Andrea, et al.
Pubblicazione: (2023)
Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced Data
di: Cortes, Corinna, et al.
Pubblicazione: (2025)
di: Cortes, Corinna, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How Well Can Preference Optimization Generalize Under Noisy Feedback?
di: Im, Shawn, et al.
Pubblicazione: (2025) -
A Unified Understanding and Evaluation of Steering Methods
di: Im, Shawn, et al.
Pubblicazione: (2025) -
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
di: Li, Wendi, et al.
Pubblicazione: (2026) -
Understanding the Learning Dynamics of Alignment with Human Feedback
di: Im, Shawn, et al.
Pubblicazione: (2024) -
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
di: Im, Shawn, et al.
Pubblicazione: (2026)