Accelerating RLHF Training with Reward Variance Increase
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zonglin, Gu, Zhexuan, Qi, Houduo, Yuan, Yancheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
by: Sheng, Jiayuan, et al.
Published: (2025)
by: Sheng, Jiayuan, et al.
Published: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
HOT: An Efficient Halpern Accelerating Algorithm for Optimal Transport Problems
by: Zhang, Guojun, et al.
Published: (2024)
by: Zhang, Guojun, et al.
Published: (2024)
Reward-Directed Score-Based Diffusion Models via q-Learning
by: Gao, Xuefeng, et al.
Published: (2024)
by: Gao, Xuefeng, et al.
Published: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
by: Yang, Shenghao, et al.
Published: (2026)
by: Yang, Shenghao, et al.
Published: (2026)
MARS: Unleashing the Power of Variance Reduction for Training Large Models
by: Yuan, Huizhuo, et al.
Published: (2024)
by: Yuan, Huizhuo, et al.
Published: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
Provable Acceleration for Diffusion Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
PID Accelerated Temporal Difference Algorithms
by: Bedaywi, Mark, et al.
Published: (2024)
by: Bedaywi, Mark, et al.
Published: (2024)
Accelerating Cutting-Plane Algorithms via Reinforcement Learning Surrogates
by: Mana, Kyle, et al.
Published: (2023)
by: Mana, Kyle, et al.
Published: (2023)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling
by: Hao, Yongchang, et al.
Published: (2026)
by: Hao, Yongchang, et al.
Published: (2026)
Training Infinitely Deep and Wide Transformers
by: Barboni, Raphaël, et al.
Published: (2026)
by: Barboni, Raphaël, et al.
Published: (2026)
A space-decoupling framework for optimization on bounded-rank matrices with orthogonally invariant constraints
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
Robust Control with Gradient Uncertainty
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Universal Approximation Theorem for Deep Q-Learning via FBSDE System
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Training Safe Neural Networks with Global SDP Bounds
by: Soletskyi, Roman, et al.
Published: (2024)
by: Soletskyi, Roman, et al.
Published: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
by: Rashidinejad, Paria, et al.
Published: (2024)
by: Rashidinejad, Paria, et al.
Published: (2024)
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
by: Zhang, Shihao, et al.
Published: (2025)
by: Zhang, Shihao, et al.
Published: (2025)
Double Momentum Method for Lower-Level Constrained Bilevel Optimization
by: Shi, Wanli, et al.
Published: (2024)
by: Shi, Wanli, et al.
Published: (2024)
The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
by: Harvey, Thomas R.
Published: (2025)
by: Harvey, Thomas R.
Published: (2025)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
by: Hrycej, Tomas, et al.
Published: (2025)
by: Hrycej, Tomas, et al.
Published: (2025)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
by: Hong, Seong-Hyun, et al.
Published: (2024)
by: Hong, Seong-Hyun, et al.
Published: (2024)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
by: Ghosh, Ipsita, et al.
Published: (2025)
by: Ghosh, Ipsita, et al.
Published: (2025)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023)
by: Farhat, Yehya, et al.
Published: (2023)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
by: Zhao, Pengxiang, et al.
Published: (2025)
by: Zhao, Pengxiang, et al.
Published: (2025)
Similar Items
-
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
by: Yang, Zonglin, et al.
Published: (2026) -
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
by: Sheng, Jiayuan, et al.
Published: (2025) -
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025) -
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025) -
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)