On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Zhirui, Tan, Vincent Y. F. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fixed-Budget Differentially Private Best Arm Identification
por: Chen, Zhirui, et al.
Publicado: (2024)
por: Chen, Zhirui, et al.
Publicado: (2024)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
por: Zhao, Qingyue, et al.
Publicado: (2026)
por: Zhao, Qingyue, et al.
Publicado: (2026)
Best Arm Identification with Possibly Biased Offline Data
por: Yang, Le, et al.
Publicado: (2025)
por: Yang, Le, et al.
Publicado: (2025)
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
por: Chen, Zhirui, et al.
Publicado: (2025)
por: Chen, Zhirui, et al.
Publicado: (2025)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
por: Chen, Fan, et al.
Publicado: (2025)
por: Chen, Fan, et al.
Publicado: (2025)
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget
por: Bian, Jie, et al.
Publicado: (2025)
por: Bian, Jie, et al.
Publicado: (2025)
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
por: Li, Yingru, et al.
Publicado: (2024)
por: Li, Yingru, et al.
Publicado: (2024)
Universal time-series forecasting with mixture predictors
por: Ryabko, Daniil
Publicado: (2020)
por: Ryabko, Daniil
Publicado: (2020)
MESSY Estimation: Maximum-Entropy based Stochastic and Symbolic densitY Estimation
por: Tohme, Tony, et al.
Publicado: (2023)
por: Tohme, Tony, et al.
Publicado: (2023)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
por: Hellström, Fredrik, et al.
Publicado: (2023)
por: Hellström, Fredrik, et al.
Publicado: (2023)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
por: Zhang, Bohan, et al.
Publicado: (2025)
por: Zhang, Bohan, et al.
Publicado: (2025)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
por: Zhang, Huiming, et al.
Publicado: (2026)
por: Zhang, Huiming, et al.
Publicado: (2026)
Greedy Sampling Is Provably Efficient for RLHF
por: Wu, Di, et al.
Publicado: (2025)
por: Wu, Di, et al.
Publicado: (2025)
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
por: Hou, Yunlong, et al.
Publicado: (2024)
por: Hou, Yunlong, et al.
Publicado: (2024)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
por: Hou, Yunlong, et al.
Publicado: (2026)
por: Hou, Yunlong, et al.
Publicado: (2026)
Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons
por: Zhu, Banghua, et al.
Publicado: (2023)
por: Zhu, Banghua, et al.
Publicado: (2023)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
por: Wu, Di, et al.
Publicado: (2026)
por: Wu, Di, et al.
Publicado: (2026)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
por: Zhao, Qingyue, et al.
Publicado: (2025)
por: Zhao, Qingyue, et al.
Publicado: (2025)
Random Multiplexing
por: Liu, Lei, et al.
Publicado: (2025)
por: Liu, Lei, et al.
Publicado: (2025)
Towards Faster Non-Asymptotic Convergence for Diffusion-Based Generative Models
por: Li, Gen, et al.
Publicado: (2023)
por: Li, Gen, et al.
Publicado: (2023)
A Fine-Grained Understanding of Uniform Convergence for Halfspaces
por: Kontorovich, Aryeh, et al.
Publicado: (2026)
por: Kontorovich, Aryeh, et al.
Publicado: (2026)
Entropy, concentration, and learning: a statistical mechanics primer
por: Balsubramani, Akshay
Publicado: (2024)
por: Balsubramani, Akshay
Publicado: (2024)
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
por: Levy, Jordan, et al.
Publicado: (2026)
por: Levy, Jordan, et al.
Publicado: (2026)
On the Separability of Information in Diffusion Models
por: Premkumar, Akhil
Publicado: (2025)
por: Premkumar, Akhil
Publicado: (2025)
Neural Estimation of Pairwise Mutual Information in Masked Discrete Sequence Models
por: Sharma, Jai, et al.
Publicado: (2026)
por: Sharma, Jai, et al.
Publicado: (2026)
Convergence of Shallow ReLU Networks on Weakly Interacting Data
por: Dana, Léo, et al.
Publicado: (2025)
por: Dana, Léo, et al.
Publicado: (2025)
Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
por: Balasubramanian, Krishnakumar
Publicado: (2026)
por: Balasubramanian, Krishnakumar
Publicado: (2026)
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
The Geometry of Knowing: From Possibilistic Ignorance to Probabilistic Certainty -- A Measure-Theoretic Framework for Epistemic Convergence
por: Jah, Moriba Kemessia
Publicado: (2026)
por: Jah, Moriba Kemessia
Publicado: (2026)
Almost Asymptotically Optimal Active Clustering Through Pairwise Observations
por: Teo, Rachel S. Y., et al.
Publicado: (2026)
por: Teo, Rachel S. Y., et al.
Publicado: (2026)
Settling the Sample Complexity of Model-Based Offline Reinforcement Learning
por: Li, Gen, et al.
Publicado: (2022)
por: Li, Gen, et al.
Publicado: (2022)
Fast Convergence of $Φ$-Divergence Along the Unadjusted Langevin Algorithm and Proximal Sampler
por: Mitra, Siddharth, et al.
Publicado: (2024)
por: Mitra, Siddharth, et al.
Publicado: (2024)
Guaranteed Recovery of Unambiguous Clusters
por: Mazooji, Kayvon, et al.
Publicado: (2025)
por: Mazooji, Kayvon, et al.
Publicado: (2025)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
por: Lu, Miao, et al.
Publicado: (2022)
por: Lu, Miao, et al.
Publicado: (2022)
Ordinary Least Squares is a Special Case of Transformer
por: Tan, Xiaojun, et al.
Publicado: (2026)
por: Tan, Xiaojun, et al.
Publicado: (2026)
Model-Based Reinforcement Learning for Offline Zero-Sum Markov Games
por: Yan, Yuling, et al.
Publicado: (2022)
por: Yan, Yuling, et al.
Publicado: (2022)
Mixed Regression via Approximate Message Passing
por: Tan, Nelvin, et al.
Publicado: (2023)
por: Tan, Nelvin, et al.
Publicado: (2023)
$L^1$ Estimation: On the Optimality of Linear Estimators
por: Barnes, Leighton P., et al.
Publicado: (2023)
por: Barnes, Leighton P., et al.
Publicado: (2023)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
por: Chen, Siyu, et al.
Publicado: (2024)
por: Chen, Siyu, et al.
Publicado: (2024)
Ejemplares similares
-
Fixed-Budget Differentially Private Best Arm Identification
por: Chen, Zhirui, et al.
Publicado: (2024) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
por: Zhao, Qingyue, et al.
Publicado: (2026) -
Best Arm Identification with Possibly Biased Offline Data
por: Yang, Le, et al.
Publicado: (2025) -
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
por: Chen, Zhirui, et al.
Publicado: (2025) -
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
por: Chen, Fan, et al.
Publicado: (2025)