Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Audrey, Zhan, Wenhao, Xie, Tengyang, Lee, Jason D., Sun, Wen, Krishnamurthy, Akshay, Foster, Dylan J. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
por: Xie, Tengyang, et al.
Publicado: (2024)
por: Xie, Tengyang, et al.
Publicado: (2024)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025)
por: Huang, Audrey, et al.
Publicado: (2025)
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
por: Rohatgi, Dhruv, et al.
Publicado: (2025)
por: Rohatgi, Dhruv, et al.
Publicado: (2025)
Scalable Online Exploration via Coverability
por: Amortila, Philip, et al.
Publicado: (2024)
por: Amortila, Philip, et al.
Publicado: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
por: Zhan, Wenhao, et al.
Publicado: (2023)
por: Zhan, Wenhao, et al.
Publicado: (2023)
Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum
por: Rajaraman, Nived, et al.
Publicado: (2026)
por: Rajaraman, Nived, et al.
Publicado: (2026)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
por: Rosset, Corby, et al.
Publicado: (2024)
por: Rosset, Corby, et al.
Publicado: (2024)
Rich-Observation Reinforcement Learning with Continuous Latent Dynamics
por: Song, Yuda, et al.
Publicado: (2024)
por: Song, Yuda, et al.
Publicado: (2024)
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
por: Tuyls, Jens, et al.
Publicado: (2025)
por: Tuyls, Jens, et al.
Publicado: (2025)
KL Penalty Control via Perturbation for Direct Preference Optimization
por: Lee, Sangkyu, et al.
Publicado: (2025)
por: Lee, Sangkyu, et al.
Publicado: (2025)
Harnessing Density Ratios for Online Reinforcement Learning
por: Amortila, Philip, et al.
Publicado: (2024)
por: Amortila, Philip, et al.
Publicado: (2024)
Forward KL Regularized Preference Optimization for Aligning Diffusion Policies
por: Shan, Zhao, et al.
Publicado: (2024)
por: Shan, Zhao, et al.
Publicado: (2024)
Can large language models explore in-context?
por: Krishnamurthy, Akshay, et al.
Publicado: (2024)
por: Krishnamurthy, Akshay, et al.
Publicado: (2024)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
por: Amortila, Philip, et al.
Publicado: (2024)
por: Amortila, Philip, et al.
Publicado: (2024)
Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases
por: Zhang, Ziyi, et al.
Publicado: (2024)
por: Zhang, Ziyi, et al.
Publicado: (2024)
Mythos
Publicado: (2023)
Publicado: (2023)
Self-Improvement in Language Models: The Sharpening Mechanism
por: Huang, Audrey, et al.
Publicado: (2024)
por: Huang, Audrey, et al.
Publicado: (2024)
The Coverage Principle: How Pre-Training Enables Post-Training
por: Chen, Fan, et al.
Publicado: (2025)
por: Chen, Fan, et al.
Publicado: (2025)
The Chi-Square Test of Distance Correlation
por: Shen, Cencheng, et al.
Publicado: (2019)
por: Shen, Cencheng, et al.
Publicado: (2019)
Square$χ$PO: Differentially Private and Robust $χ^2$-Preference Optimization in Offline Direct Alignment
por: Zhou, Xingyu, et al.
Publicado: (2025)
por: Zhou, Xingyu, et al.
Publicado: (2025)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
por: Liu, Zhihan, et al.
Publicado: (2024)
por: Liu, Zhihan, et al.
Publicado: (2024)
Managerial Overoptimism and Discretionary Disclosure
por: Nikolaj Niebuhr Lambertsen, et al.
Publicado: (2026)
por: Nikolaj Niebuhr Lambertsen, et al.
Publicado: (2026)
Mythos Enigma
por: Landwehr, Dominik
Publicado: (2020)
por: Landwehr, Dominik
Publicado: (2020)
Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
por: Du, Jie, et al.
Publicado: (2025)
por: Du, Jie, et al.
Publicado: (2025)
A Unifying View of Coverage in Linear Off-Policy Evaluation
por: Amortila, Philip, et al.
Publicado: (2026)
por: Amortila, Philip, et al.
Publicado: (2026)
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
por: Lu, Junru, et al.
Publicado: (2024)
por: Lu, Junru, et al.
Publicado: (2024)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
por: Chen, Ziyi, et al.
Publicado: (2025)
por: Chen, Ziyi, et al.
Publicado: (2025)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
por: Liu, Kezhao, et al.
Publicado: (2025)
por: Liu, Kezhao, et al.
Publicado: (2025)
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
por: Liu, Jilong, et al.
Publicado: (2026)
por: Liu, Jilong, et al.
Publicado: (2026)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
por: Ji, Xiang, et al.
Publicado: (2024)
por: Ji, Xiang, et al.
Publicado: (2024)
Inverse Reinforcement Learning using Revealed Preferences and Passive Stochastic Optimization
por: Krishnamurthy, Vikram
Publicado: (2025)
por: Krishnamurthy, Vikram
Publicado: (2025)
Enhancing Model Fit Evaluation in SEM: Practical Tips for Optimizing Chi-Square Tests
por: Zheng, Bang Quan, et al.
Publicado: (2023)
por: Zheng, Bang Quan, et al.
Publicado: (2023)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
por: Kim, Sunghwan, et al.
Publicado: (2025)
por: Kim, Sunghwan, et al.
Publicado: (2025)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
por: Jiang, Nan, et al.
Publicado: (2025)
por: Jiang, Nan, et al.
Publicado: (2025)
Understanding Behavior Cloning with Action Quantization
por: Cao, Haoqun, et al.
Publicado: (2026)
por: Cao, Haoqun, et al.
Publicado: (2026)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
por: Yuan, Yurun, et al.
Publicado: (2026)
por: Yuan, Yurun, et al.
Publicado: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
por: Yuan, Yurun, et al.
Publicado: (2025)
por: Yuan, Yurun, et al.
Publicado: (2025)
Chi-Square Wavelet Graph Neural Networks for Heterogeneous Graph Anomaly Detection
por: Li, Xiping, et al.
Publicado: (2025)
por: Li, Xiping, et al.
Publicado: (2025)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
por: Wang, Haoxiang, et al.
Publicado: (2024)
por: Wang, Haoxiang, et al.
Publicado: (2024)
Ejemplares similares
-
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
por: Xie, Tengyang, et al.
Publicado: (2024) -
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025) -
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
por: Rohatgi, Dhruv, et al.
Publicado: (2025) -
Scalable Online Exploration via Coverability
por: Amortila, Philip, et al.
Publicado: (2024) -
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
por: Rafailov, Rafael, et al.
Publicado: (2024)