Towards a Theoretical Understanding to the Generalization of RLHF
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zhaochun, Yi, Mingyang, Wang, Yue, Cui, Shisheng, Liu, Yong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
di: Li, Zhaochun, et al.
Pubblicazione: (2026)
di: Li, Zhaochun, et al.
Pubblicazione: (2026)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
di: Zhang, Huiming, et al.
Pubblicazione: (2026)
di: Zhang, Huiming, et al.
Pubblicazione: (2026)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
di: Ouyang, Sheng, et al.
Pubblicazione: (2025)
di: Ouyang, Sheng, et al.
Pubblicazione: (2025)
Understanding and Alleviating Memory Consumption in RLHF for LLMs
di: Zhou, Jin, et al.
Pubblicazione: (2024)
di: Zhou, Jin, et al.
Pubblicazione: (2024)
Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model
di: Yi, Mingyang, et al.
Pubblicazione: (2024)
di: Yi, Mingyang, et al.
Pubblicazione: (2024)
Beyond RLHF: A Unified Theoretical Framework of Alignment
di: Yun, Jihun, et al.
Pubblicazione: (2025)
di: Yun, Jihun, et al.
Pubblicazione: (2025)
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
di: Park, Chanwoo, et al.
Pubblicazione: (2024)
di: Park, Chanwoo, et al.
Pubblicazione: (2024)
Towards Theoretical Understandings of Self-Consuming Generative Models
di: Fu, Shi, et al.
Pubblicazione: (2024)
di: Fu, Shi, et al.
Pubblicazione: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
di: Mao, Liyuan, et al.
Pubblicazione: (2025)
di: Mao, Liyuan, et al.
Pubblicazione: (2025)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
di: Gan, Zeyu, et al.
Pubblicazione: (2024)
di: Gan, Zeyu, et al.
Pubblicazione: (2024)
Mitigating the Alignment Tax of RLHF
di: Lin, Yong, et al.
Pubblicazione: (2023)
di: Lin, Yong, et al.
Pubblicazione: (2023)
A Theoretical Framework for Partially Observed Reward-States in RLHF
di: Kausik, Chinmaya, et al.
Pubblicazione: (2024)
di: Kausik, Chinmaya, et al.
Pubblicazione: (2024)
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
di: Xiao, Jiancong, et al.
Pubblicazione: (2025)
di: Xiao, Jiancong, et al.
Pubblicazione: (2025)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
di: Miao, Yuchun, et al.
Pubblicazione: (2025)
di: Miao, Yuchun, et al.
Pubblicazione: (2025)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
Stability and Sharper Risk Bounds with Convergence Rate $\tilde{O}(1/n^2)$
di: Zhu, Bowei, et al.
Pubblicazione: (2024)
di: Zhu, Bowei, et al.
Pubblicazione: (2024)
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
di: He, Pengfei, et al.
Pubblicazione: (2024)
di: He, Pengfei, et al.
Pubblicazione: (2024)
Unifying Stable Optimization and Reference Regularization in RLHF
di: He, Li, et al.
Pubblicazione: (2026)
di: He, Li, et al.
Pubblicazione: (2026)
Towards Reliable Alignment: Uncertainty-aware RLHF
di: Banerjee, Debangshu, et al.
Pubblicazione: (2024)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2024)
Towards Federated RLHF with Aggregated Client Preference for LLMs
di: Wu, Feijie, et al.
Pubblicazione: (2024)
di: Wu, Feijie, et al.
Pubblicazione: (2024)
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
di: Xu, Wenyuan, et al.
Pubblicazione: (2025)
di: Xu, Wenyuan, et al.
Pubblicazione: (2025)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
di: Kirk, Robert, et al.
Pubblicazione: (2023)
di: Kirk, Robert, et al.
Pubblicazione: (2023)
General Exploratory Bonus for Optimistic Exploration in RLHF
di: Li, Wendi, et al.
Pubblicazione: (2025)
di: Li, Wendi, et al.
Pubblicazione: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
di: Siththaranjan, Anand, et al.
Pubblicazione: (2023)
di: Siththaranjan, Anand, et al.
Pubblicazione: (2023)
Thompson Sampling in Online RLHF with General Function Approximation
di: Feng, Songtao, et al.
Pubblicazione: (2025)
di: Feng, Songtao, et al.
Pubblicazione: (2025)
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
di: Tang, Huayi, et al.
Pubblicazione: (2023)
di: Tang, Huayi, et al.
Pubblicazione: (2023)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
On a Connection Between Imitation Learning and RLHF
di: Xiao, Teng, et al.
Pubblicazione: (2025)
di: Xiao, Teng, et al.
Pubblicazione: (2025)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
di: Hu, Jian, et al.
Pubblicazione: (2024)
di: Hu, Jian, et al.
Pubblicazione: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
di: Sheng, Jiayuan, et al.
Pubblicazione: (2025)
di: Sheng, Jiayuan, et al.
Pubblicazione: (2025)
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
di: Mukherjee, Arpan, et al.
Pubblicazione: (2025)
di: Mukherjee, Arpan, et al.
Pubblicazione: (2025)
Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
di: Ren, Ruifeng, et al.
Pubblicazione: (2023)
di: Ren, Ruifeng, et al.
Pubblicazione: (2023)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
di: Zhou, Xingyu, et al.
Pubblicazione: (2025)
di: Zhou, Xingyu, et al.
Pubblicazione: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
di: Shi, Ruizhe, et al.
Pubblicazione: (2025)
di: Shi, Ruizhe, et al.
Pubblicazione: (2025)
An Inclusive Theoretical Framework of Robust Supervised Contrastive Loss against Label Noise
di: Cui, Jingyi, et al.
Pubblicazione: (2025)
di: Cui, Jingyi, et al.
Pubblicazione: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024)
di: Dong, Hanze, et al.
Pubblicazione: (2024)
Continuous-time Riemannian SGD and SVRG Flows on Wasserstein Probabilistic Space
di: Yi, Mingyang, et al.
Pubblicazione: (2024)
di: Yi, Mingyang, et al.
Pubblicazione: (2024)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
di: Liu, Renpu, et al.
Pubblicazione: (2025)
di: Liu, Renpu, et al.
Pubblicazione: (2025)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
di: Du, Yihan, et al.
Pubblicazione: (2024)
di: Du, Yihan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
di: Li, Zhaochun, et al.
Pubblicazione: (2026) -
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
di: Zhang, Huiming, et al.
Pubblicazione: (2026) -
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
di: Ouyang, Sheng, et al.
Pubblicazione: (2025) -
Understanding and Alleviating Memory Consumption in RLHF for LLMs
di: Zhou, Jin, et al.
Pubblicazione: (2024) -
Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model
di: Yi, Mingyang, et al.
Pubblicazione: (2024)