An Approximate Ascent Approach To Prove Convergence of PPO
Fuente:
arXiv
Salvato in:
| Autori principali: | Doering, Leif, Schmidt, Daniel, Melcher, Moritz, Kassing, Sebastian, Wille, Benedikt, Aach, Tilman, Weissmann, Simon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Role of Target Update Frequencies in Q-Learning
di: Weissmann, Simon, et al.
Pubblicazione: (2026)
di: Weissmann, Simon, et al.
Pubblicazione: (2026)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
di: Kassing, Sebastian, et al.
Pubblicazione: (2025)
di: Kassing, Sebastian, et al.
Pubblicazione: (2025)
Polyak's Heavy Ball Method Achieves Accelerated Local Rate of Convergence under Polyak-Lojasiewicz Inequality
di: Kassing, Sebastian, et al.
Pubblicazione: (2024)
di: Kassing, Sebastian, et al.
Pubblicazione: (2024)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
di: Klein, Sara, et al.
Pubblicazione: (2023)
di: Klein, Sara, et al.
Pubblicazione: (2023)
An Improved Last-Iterate Convergence Rate for Anchored Gradient Descent Ascent
di: Surina, Anja, et al.
Pubblicazione: (2026)
di: Surina, Anja, et al.
Pubblicazione: (2026)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
di: Zhang, Qixin, et al.
Pubblicazione: (2024)
di: Zhang, Qixin, et al.
Pubblicazione: (2024)
Almost sure convergence rates of stochastic gradient methods under gradient domination
di: Weissmann, Simon, et al.
Pubblicazione: (2024)
di: Weissmann, Simon, et al.
Pubblicazione: (2024)
ADDQ: Adaptive Distributional Double Q-Learning
di: Döring, Leif, et al.
Pubblicazione: (2025)
di: Döring, Leif, et al.
Pubblicazione: (2025)
A Single-Loop Gradient Descent and Perturbed Ascent Algorithm for Nonconvex Functional Constrained Optimization
di: Lu, Songtao
Pubblicazione: (2022)
di: Lu, Songtao
Pubblicazione: (2022)
Non-Asymptotic Global Convergence of PPO-Clip
di: Liu, Yin, et al.
Pubblicazione: (2025)
di: Liu, Yin, et al.
Pubblicazione: (2025)
Structure Matters: Dynamic Policy Gradient
di: Klein, Sara, et al.
Pubblicazione: (2024)
di: Klein, Sara, et al.
Pubblicazione: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
di: Gao, Yihang, et al.
Pubblicazione: (2024)
di: Gao, Yihang, et al.
Pubblicazione: (2024)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
di: Ganesh, Swetha, et al.
Pubblicazione: (2024)
di: Ganesh, Swetha, et al.
Pubblicazione: (2024)
RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection
di: Hajieghrary, Hadi, et al.
Pubblicazione: (2026)
di: Hajieghrary, Hadi, et al.
Pubblicazione: (2026)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
di: Schotthöfer, Steffen, et al.
Pubblicazione: (2024)
di: Schotthöfer, Steffen, et al.
Pubblicazione: (2024)
Convergence of Some Convex Message Passing Algorithms to a Fixed Point
di: Voracek, Vaclav, et al.
Pubblicazione: (2024)
di: Voracek, Vaclav, et al.
Pubblicazione: (2024)
Sobolev Gradient Ascent for Optimal Transport: Barycenter Optimization and Convergence Analysis
di: Kim, Kaheon, et al.
Pubblicazione: (2025)
di: Kim, Kaheon, et al.
Pubblicazione: (2025)
On the Convergence of Overparameterized Problems: Inherent Properties of the Compositional Structure of Neural Networks
di: de Oliveira, Arthur Castello Branco, et al.
Pubblicazione: (2025)
di: de Oliveira, Arthur Castello Branco, et al.
Pubblicazione: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
di: Xiao, Minheng, et al.
Pubblicazione: (2024)
di: Xiao, Minheng, et al.
Pubblicazione: (2024)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
di: Ma, Shaocong, et al.
Pubblicazione: (2026)
di: Ma, Shaocong, et al.
Pubblicazione: (2026)
Global Convergence of Multiplicative Updates for the Matrix Mechanism: A Collaborative Proof with Gemini 3
di: Rush, Keith
Pubblicazione: (2026)
di: Rush, Keith
Pubblicazione: (2026)
A Methodology Establishing Linear Convergence of Adaptive Gradient Methods under PL Inequality
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2024)
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
di: Wagner, Dominik, et al.
Pubblicazione: (2024)
ODE approximation for the Adam algorithm: General and overparametrized setting
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
di: Dereich, Steffen, et al.
Pubblicazione: (2025)
Fitted Q-Iteration via Max-Plus-Linear Approximation
di: Liu, Y., et al.
Pubblicazione: (2024)
di: Liu, Y., et al.
Pubblicazione: (2024)
Ginger: An Efficient Curvature Approximation with Linear Complexity for General Neural Networks
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian Noise
di: Blaser, Ethan, et al.
Pubblicazione: (2024)
di: Blaser, Ethan, et al.
Pubblicazione: (2024)
Universal Approximation Theorem for Deep Q-Learning via FBSDE System
di: Qi, Qian
Pubblicazione: (2025)
di: Qi, Qian
Pubblicazione: (2025)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
di: Xu, Conglong, et al.
Pubblicazione: (2025)
di: Xu, Conglong, et al.
Pubblicazione: (2025)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
di: Ding, Dongsheng, et al.
Pubblicazione: (2022)
di: Ding, Dongsheng, et al.
Pubblicazione: (2022)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
di: Chen, Zixiang, et al.
Pubblicazione: (2025)
di: Chen, Zixiang, et al.
Pubblicazione: (2025)
Differentiable Optimization for Deep Learning-Enhanced DC Approximation of AC Optimal Power Flow
di: Rosemberg, Andrew, et al.
Pubblicazione: (2025)
di: Rosemberg, Andrew, et al.
Pubblicazione: (2025)
Achieving Tighter Finite-Time Rates for Heterogeneous Federated Stochastic Approximation under Markovian Sampling
di: Zhu, Feng, et al.
Pubblicazione: (2025)
di: Zhu, Feng, et al.
Pubblicazione: (2025)
DASA: Delay-Adaptive Multi-Agent Stochastic Approximation
di: Fabbro, Nicolò Dal, et al.
Pubblicazione: (2024)
di: Fabbro, Nicolò Dal, et al.
Pubblicazione: (2024)
Dual Optimistic Ascent (PI Control) is the Augmented Lagrangian Method in Disguise
di: Ramirez, Juan, et al.
Pubblicazione: (2025)
di: Ramirez, Juan, et al.
Pubblicazione: (2025)
Seeing Through Risk: A Symbolic Approximation of Prospect Theory
di: Yousaf, Ali Arslan, et al.
Pubblicazione: (2025)
di: Yousaf, Ali Arslan, et al.
Pubblicazione: (2025)
Negative Stepsizes Make Gradient-Descent-Ascent Converge
di: Shugart, Henry, et al.
Pubblicazione: (2025)
di: Shugart, Henry, et al.
Pubblicazione: (2025)
Differentiable Nonlinear Model Predictive Control
di: Frey, Jonathan, et al.
Pubblicazione: (2025)
di: Frey, Jonathan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Role of Target Update Frequencies in Q-Learning
di: Weissmann, Simon, et al.
Pubblicazione: (2026) -
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
di: Kassing, Sebastian, et al.
Pubblicazione: (2025) -
Polyak's Heavy Ball Method Achieves Accelerated Local Rate of Convergence under Polyak-Lojasiewicz Inequality
di: Kassing, Sebastian, et al.
Pubblicazione: (2024) -
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
di: Klein, Sara, et al.
Pubblicazione: (2023) -
An Improved Last-Iterate Convergence Rate for Anchored Gradient Descent Ascent
di: Surina, Anja, et al.
Pubblicazione: (2026)