On the optimization dynamics of RLVR: Gradient gap and step size thresholds
Fuente:
arXiv
Saved in:
| Main Authors: | Suk, Joe, Duan, Yaqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Taming "data-hungry" reinforcement learning? Stability in continuous state-action spaces
by: Duan, Yaqi, et al.
Published: (2024)
by: Duan, Yaqi, et al.
Published: (2024)
One if by Land, Two if by Sea, Three if by Four Seas, and More to Come -- Values of Perception, Prediction, Communication, and Common Sense in Decision Making
by: Xu, Aolin
Published: (2025)
by: Xu, Aolin
Published: (2025)
Capabilities and Fundamental Limits of Latent Chain-of-Thought
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
by: Shen, Wei, et al.
Published: (2023)
by: Shen, Wei, et al.
Published: (2023)
Queueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers
by: Ozbas, Emre, et al.
Published: (2026)
by: Ozbas, Emre, et al.
Published: (2026)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Delightful Policy Gradient
by: Osband, Ian
Published: (2026)
by: Osband, Ian
Published: (2026)
Robust Control with Gradient Uncertainty
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Delightful Distributed Policy Gradient
by: Osband, Ian
Published: (2026)
by: Osband, Ian
Published: (2026)
Gradient Descent Efficiency Index
by: Dhingra, Aviral
Published: (2024)
by: Dhingra, Aviral
Published: (2024)
Exploring Applications of State Space Models and Advanced Training Techniques in Sequential Recommendations: A Comparative Study on Efficiency and Performance
by: Obozov, Mark, et al.
Published: (2024)
by: Obozov, Mark, et al.
Published: (2024)
Long-time dynamics and universality of nonconvex gradient descent
by: Han, Qiyang
Published: (2025)
by: Han, Qiyang
Published: (2025)
Gradient descent inference in empirical risk minimization
by: Han, Qiyang, et al.
Published: (2024)
by: Han, Qiyang, et al.
Published: (2024)
Unbiased Gradient Low-Rank Projection
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Tutorial on amortized optimization
by: Amos, Brandon
Published: (2022)
by: Amos, Brandon
Published: (2022)
Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method
by: Bright-Thonney, Samuel, et al.
Published: (2026)
by: Bright-Thonney, Samuel, et al.
Published: (2026)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
by: Hennick, Max, et al.
Published: (2025)
by: Hennick, Max, et al.
Published: (2025)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
by: Zhang, Qixin, et al.
Published: (2024)
by: Zhang, Qixin, et al.
Published: (2024)
Riemann Sum Optimization for Accurate Integrated Gradients Computation
by: Swain, Swadesh, et al.
Published: (2024)
by: Swain, Swadesh, et al.
Published: (2024)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
by: Arda, Enes, et al.
Published: (2026)
by: Arda, Enes, et al.
Published: (2026)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
by: Gao, Yihang, et al.
Published: (2024)
by: Gao, Yihang, et al.
Published: (2024)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
by: Hu, Zeou, et al.
Published: (2026)
by: Hu, Zeou, et al.
Published: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
by: Xu, Conglong, et al.
Published: (2025)
by: Xu, Conglong, et al.
Published: (2025)
On a Gradient Approach to Chebyshev Center Problems with Applications to Function Learning
by: Raghuvanshi, Abhinav, et al.
Published: (2026)
by: Raghuvanshi, Abhinav, et al.
Published: (2026)
Towards Efficient Risk-Sensitive Policy Gradient: An Iteration Complexity Analysis
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Fuzzy hyperparameters update in a second order optimization
by: Bensadok, Abdelaziz, et al.
Published: (2024)
by: Bensadok, Abdelaziz, et al.
Published: (2024)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
Multi-beam Beamforming in RIS-aided MIMO Subject to Reradiation Mask Constraints -- Optimization and Machine Learning Design
by: Wang, Shumin, et al.
Published: (2025)
by: Wang, Shumin, et al.
Published: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
by: Alvo, Matias, et al.
Published: (2026)
by: Alvo, Matias, et al.
Published: (2026)
Similar Items
-
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025) -
Taming "data-hungry" reinforcement learning? Stability in continuous state-action spaces
by: Duan, Yaqi, et al.
Published: (2024) -
One if by Land, Two if by Sea, Three if by Four Seas, and More to Come -- Values of Perception, Prediction, Communication, and Common Sense in Decision Making
by: Xu, Aolin
Published: (2025) -
Capabilities and Fundamental Limits of Latent Chain-of-Thought
by: Zou, Jiaxuan, et al.
Published: (2026) -
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)