Greedy Sampling Is Provably Efficient for RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Di, Shi, Chengshuai, Yang, Jing, Shen, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Cost-Aware Optimal Pairwise Pure Exploration
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
by: Cai, Changxiao, et al.
Published: (2026)
by: Cai, Changxiao, et al.
Published: (2026)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
by: Li, Donghao, et al.
Published: (2026)
by: Li, Donghao, et al.
Published: (2026)
Harnessing the Power of Federated Learning in Federated Contextual Bandits
by: Shi, Chengshuai, et al.
Published: (2023)
by: Shi, Chengshuai, et al.
Published: (2023)
Provable Privacy Advantages of Decentralized Federated Learning via Distributed Optimization
by: Yu, Wenrui, et al.
Published: (2024)
by: Yu, Wenrui, et al.
Published: (2024)
Efficient Prompt Optimization Through the Lens of Best Arm Identification
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
Digital Over-the-Air Federated Learning in Multi-Antenna Systems
by: Wang, Sihua, et al.
Published: (2023)
by: Wang, Sihua, et al.
Published: (2023)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
by: Liu, Renpu, et al.
Published: (2025)
by: Liu, Renpu, et al.
Published: (2025)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
Order-Optimal Sample Complexity of Rectified Flows
by: Sahoo, Hari Krishna, et al.
Published: (2026)
by: Sahoo, Hari Krishna, et al.
Published: (2026)
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
by: Kaiser, Johannes, et al.
Published: (2025)
by: Kaiser, Johannes, et al.
Published: (2025)
Integrated Sensing-Communication-Computation for Edge Artificial Intelligence
by: Wen, Dingzhu, et al.
Published: (2023)
by: Wen, Dingzhu, et al.
Published: (2023)
An Information-Theoretic Criterion for Efficient Data Synthesis
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
by: Wang, Yikun, et al.
Published: (2026)
by: Wang, Yikun, et al.
Published: (2026)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
by: Bhattacharjee, Payel, et al.
Published: (2025)
by: Bhattacharjee, Payel, et al.
Published: (2025)
Energy-Efficient Edge Learning via Joint Data Deepening-and-Prefetching
by: Kook, Sujin, et al.
Published: (2024)
by: Kook, Sujin, et al.
Published: (2024)
Dual-Domain Deep Learning-Assisted NOMA-CSK Systems for Secure and Efficient Vehicular Communications
by: Huang, Tingting, et al.
Published: (2025)
by: Huang, Tingting, et al.
Published: (2025)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
by: Yang, Zhiqin, et al.
Published: (2026)
by: Yang, Zhiqin, et al.
Published: (2026)
Knowledge Graph-Based Explainable and Generalized Zero-Shot Semantic Communications
by: Zhang, Zhaoyu, et al.
Published: (2025)
by: Zhang, Zhaoyu, et al.
Published: (2025)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
A Wireless Foundation Model for Multi-Task Prediction
by: Sheng, Yucheng, et al.
Published: (2025)
by: Sheng, Yucheng, et al.
Published: (2025)
Learning Unknown Intervention Targets in Structural Causal Models from Heterogeneous Data
by: Yang, Yuqin, et al.
Published: (2023)
by: Yang, Yuqin, et al.
Published: (2023)
General Information Metrics for Improving AI Model Training Efficiency
by: Xu, Jianfeng, et al.
Published: (2025)
by: Xu, Jianfeng, et al.
Published: (2025)
Semi-supervised Batch Learning From Logged Data
by: Aminian, Gholamali, et al.
Published: (2022)
by: Aminian, Gholamali, et al.
Published: (2022)
Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter Selection
by: Zecchin, Matteo, et al.
Published: (2024)
by: Zecchin, Matteo, et al.
Published: (2024)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
by: Yang, Jiaming, et al.
Published: (2026)
by: Yang, Jiaming, et al.
Published: (2026)
Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning
by: Wang, Zhongwei, et al.
Published: (2025)
by: Wang, Zhongwei, et al.
Published: (2025)
Two Birds with One Stone: Multi-Task Semantic Communications Systems over Relay Channel
by: Cao, Yujie, et al.
Published: (2024)
by: Cao, Yujie, et al.
Published: (2024)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
Foundational theories of hesitant fuzzy sets and families of hesitant fuzzy sets
by: Lu, Shizhan, et al.
Published: (2023)
by: Lu, Shizhan, et al.
Published: (2023)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory
by: Dong, Songwei, et al.
Published: (2026)
by: Dong, Songwei, et al.
Published: (2026)
Similar Items
-
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
by: Wu, Di, et al.
Published: (2026) -
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024) -
Cost-Aware Optimal Pairwise Pure Exploration
by: Wu, Di, et al.
Published: (2025) -
Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
by: Cai, Changxiao, et al.
Published: (2026) -
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024)