Beyond RLHF: A Unified Theoretical Framework of Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yun, Jihun, Kim, Juno, Park, Jongho, Kim, Junhyuck, Ryu, Jongha Jon, Cho, Jaewoong, Jun, Kwang-Sung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Coverage Improvement and Fast Convergence of On-policy Preference Learning
von: Kim, Juno, et al.
Veröffentlicht: (2026)
von: Kim, Juno, et al.
Veröffentlicht: (2026)
Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2025)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2025)
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
Regularized Online RLHF with Generalized Bilinear Preferences
von: Lee, Junghyun, et al.
Veröffentlicht: (2026)
von: Lee, Junghyun, et al.
Veröffentlicht: (2026)
Task Diversity Shortens the ICL Plateau
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
von: Lee, Junghyun, et al.
Veröffentlicht: (2024)
von: Lee, Junghyun, et al.
Veröffentlicht: (2024)
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024)
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
von: Zhou, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhou, Xingyu, et al.
Veröffentlicht: (2025)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)
Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale
von: Park, Sihwan, et al.
Veröffentlicht: (2025)
von: Park, Sihwan, et al.
Veröffentlicht: (2025)
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
von: Zhao, Yao, et al.
Veröffentlicht: (2026)
von: Zhao, Yao, et al.
Veröffentlicht: (2026)
Understanding active learning of molecular docking and its applications
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2024)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024)
von: Park, Jongho, et al.
Veröffentlicht: (2024)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
von: Park, Inkyu, et al.
Veröffentlicht: (2023)
von: Park, Inkyu, et al.
Veröffentlicht: (2023)
A Theoretical Framework for Partially Observed Reward-States in RLHF
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
SLADE: Detecting Dynamic Anomalies in Edge Streams without Labels via Self-Supervised Learning
von: Lee, Jongha, et al.
Veröffentlicht: (2024)
von: Lee, Jongha, et al.
Veröffentlicht: (2024)
Mitigating the Alignment Tax of RLHF
von: Lin, Yong, et al.
Veröffentlicht: (2023)
von: Lin, Yong, et al.
Veröffentlicht: (2023)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
SeedFlood: A Step Toward Scalable Decentralized Training of LLMs
von: Kim, Jihun, et al.
Veröffentlicht: (2026)
von: Kim, Jihun, et al.
Veröffentlicht: (2026)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyun, et al.
Veröffentlicht: (2025)
Fast and Accurate Neural Rendering Using Semi-Gradients
von: Cho, In-Young, et al.
Veröffentlicht: (2024)
von: Cho, In-Young, et al.
Veröffentlicht: (2024)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
Mirror Mean-Field Langevin Dynamics
von: Gu, Anming, et al.
Veröffentlicht: (2025)
von: Gu, Anming, et al.
Veröffentlicht: (2025)
CCNETS: A Modular Causal Learning Framework for Pattern Recognition in Imbalanced Datasets
von: Park, Hanbeot, et al.
Veröffentlicht: (2024)
von: Park, Hanbeot, et al.
Veröffentlicht: (2024)
Towards a Theoretical Understanding to the Generalization of RLHF
von: Li, Zhaochun, et al.
Veröffentlicht: (2026)
von: Li, Zhaochun, et al.
Veröffentlicht: (2026)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
von: Cho, Taehyun, et al.
Veröffentlicht: (2025)
von: Cho, Taehyun, et al.
Veröffentlicht: (2025)
When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE
von: Erol, Melihcan, et al.
Veröffentlicht: (2026)
von: Erol, Melihcan, et al.
Veröffentlicht: (2026)
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
von: Xu, Wenyuan, et al.
Veröffentlicht: (2025)
von: Xu, Wenyuan, et al.
Veröffentlicht: (2025)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Explaining and Preventing Alignment Collapse in Iterative RLHF
von: Gauthier, Etienne, et al.
Veröffentlicht: (2026)
von: Gauthier, Etienne, et al.
Veröffentlicht: (2026)
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
von: Li, Yinan, et al.
Veröffentlicht: (2025)
von: Li, Yinan, et al.
Veröffentlicht: (2025)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
Unifying Stable Optimization and Reference Regularization in RLHF
von: He, Li, et al.
Veröffentlicht: (2026)
von: He, Li, et al.
Veröffentlicht: (2026)
Data-driven development of cycle prediction models for lithium metal batteries using multi modal mining
von: Lee, Jaewoong, et al.
Veröffentlicht: (2024)
von: Lee, Jaewoong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024) -
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026) -
Coverage Improvement and Fast Convergence of On-policy Preference Learning
von: Kim, Juno, et al.
Veröffentlicht: (2026) -
Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2025) -
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)