$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Di, Shi, Chengshuai, Yang, Jing, Shen, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Cost-Aware Optimal Pairwise Pure Exploration
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Harnessing the Power of Federated Learning in Federated Contextual Bandits
by: Shi, Chengshuai, et al.
Published: (2023)
by: Shi, Chengshuai, et al.
Published: (2023)
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
by: Li, Donghao, et al.
Published: (2026)
by: Li, Donghao, et al.
Published: (2026)
Digital Over-the-Air Federated Learning in Multi-Antenna Systems
by: Wang, Sihua, et al.
Published: (2023)
by: Wang, Sihua, et al.
Published: (2023)
Enhancing Explainability of Graph Neural Networks Through Conceptual and Structural Analyses and Their Extensions
by: Bui, Tien Cuong
Published: (2025)
by: Bui, Tien Cuong
Published: (2025)
Machine Unlearning via Information Theoretic Regularization
by: Xu, Shizhou, et al.
Published: (2025)
by: Xu, Shizhou, et al.
Published: (2025)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
by: Yang, Jiaming, et al.
Published: (2026)
by: Yang, Jiaming, et al.
Published: (2026)
Equivalence of the Empirical Risk Minimization to Regularization on the Family of f-Divergences
by: Daunas, Francisco, et al.
Published: (2024)
by: Daunas, Francisco, et al.
Published: (2024)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
by: Liu, Renpu, et al.
Published: (2025)
by: Liu, Renpu, et al.
Published: (2025)
Two Birds with One Stone: Multi-Task Semantic Communications Systems over Relay Channel
by: Cao, Yujie, et al.
Published: (2024)
by: Cao, Yujie, et al.
Published: (2024)
Order-Optimal Sample Complexity of Rectified Flows
by: Sahoo, Hari Krishna, et al.
Published: (2026)
by: Sahoo, Hari Krishna, et al.
Published: (2026)
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
by: Kaiser, Johannes, et al.
Published: (2025)
by: Kaiser, Johannes, et al.
Published: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
Integrated Sensing-Communication-Computation for Edge Artificial Intelligence
by: Wen, Dingzhu, et al.
Published: (2023)
by: Wen, Dingzhu, et al.
Published: (2023)
Efficient Prompt Optimization Through the Lens of Best Arm Identification
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Towards Cohesion-Fairness Harmony: Contrastive Regularization in Individual Fair Graph Clustering
by: Ghodsi, Siamak, et al.
Published: (2024)
by: Ghodsi, Siamak, et al.
Published: (2024)
A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
by: Wang, Yikun, et al.
Published: (2026)
by: Wang, Yikun, et al.
Published: (2026)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
Knowledge Graph-Based Explainable and Generalized Zero-Shot Semantic Communications
by: Zhang, Zhaoyu, et al.
Published: (2025)
by: Zhang, Zhaoyu, et al.
Published: (2025)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
A Wireless Foundation Model for Multi-Task Prediction
by: Sheng, Yucheng, et al.
Published: (2025)
by: Sheng, Yucheng, et al.
Published: (2025)
Offline and Online KL-Regularized RLHF under Differential Privacy
by: Wu, Yulian, et al.
Published: (2025)
by: Wu, Yulian, et al.
Published: (2025)
Learning Unknown Intervention Targets in Structural Causal Models from Heterogeneous Data
by: Yang, Yuqin, et al.
Published: (2023)
by: Yang, Yuqin, et al.
Published: (2023)
General Information Metrics for Improving AI Model Training Efficiency
by: Xu, Jianfeng, et al.
Published: (2025)
by: Xu, Jianfeng, et al.
Published: (2025)
Robust Semi-supervised Learning via $f$-Divergence and $α$-Rényi Divergence
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
Semi-supervised Batch Learning From Logged Data
by: Aminian, Gholamali, et al.
Published: (2022)
by: Aminian, Gholamali, et al.
Published: (2022)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
by: Garg, Sahil, et al.
Published: (2024)
by: Garg, Sahil, et al.
Published: (2024)
Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning
by: Wang, Zhongwei, et al.
Published: (2025)
by: Wang, Zhongwei, et al.
Published: (2025)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
Foundational theories of hesitant fuzzy sets and families of hesitant fuzzy sets
by: Lu, Shizhan, et al.
Published: (2023)
by: Lu, Shizhan, et al.
Published: (2023)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory
by: Dong, Songwei, et al.
Published: (2026)
by: Dong, Songwei, et al.
Published: (2026)
Statistical Channel Fingerprint Construction for Massive MIMO: A Unified Tensor Learning Framework
by: Jin, Zhenzhou, et al.
Published: (2026)
by: Jin, Zhenzhou, et al.
Published: (2026)
One Model, Two Markets: Bid-Aware Generative Recommendation
by: Jiang, Yanchen, et al.
Published: (2026)
by: Jiang, Yanchen, et al.
Published: (2026)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
by: Shen, Wei, et al.
Published: (2024)
by: Shen, Wei, et al.
Published: (2024)
Similar Items
-
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025) -
Cost-Aware Optimal Pairwise Pure Exploration
by: Wu, Di, et al.
Published: (2025) -
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024) -
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024) -
Harnessing the Power of Federated Learning in Federated Contextual Bandits
by: Shi, Chengshuai, et al.
Published: (2023)