Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
Fuente:
arXiv
Saved in:
| Main Authors: | Gaur, Mudit, Bedi, Amrit Singh, Wang, Di, Aggarwal, Vaneet |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On The Global Convergence Of Online RLHF With Neural Parametrization
by: Gaur, Mudit, et al.
Published: (2024)
by: Gaur, Mudit, et al.
Published: (2024)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)
by: Bai, Qinbo, et al.
Published: (2022)
Discrete State Diffusion Models: A Sample Complexity Perspective
by: Srikanth, Aadithya, et al.
Published: (2025)
by: Srikanth, Aadithya, et al.
Published: (2025)
Order-Optimal Sample Complexity of Rectified Flows
by: Sahoo, Hari Krishna, et al.
Published: (2026)
by: Sahoo, Hari Krishna, et al.
Published: (2026)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
BalancedDPO: Adaptive Multi-Metric Alignment
by: Tamboli, Dipesh, et al.
Published: (2025)
by: Tamboli, Dipesh, et al.
Published: (2025)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025)
by: Trivedi, Prashant, et al.
Published: (2025)
Sample Complexity Analysis for Constrained Bilevel Reinforcement Learning
by: Saxena, Naman, et al.
Published: (2026)
by: Saxena, Naman, et al.
Published: (2026)
Learning Multi-Robot Coordination through Locality-Based Factorized Multi-Agent Actor-Critic Algorithm
by: Shek, Chak Lam, et al.
Published: (2025)
by: Shek, Chak Lam, et al.
Published: (2025)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Actor-Critics Can Achieve Optimal Sample Efficiency
by: Tan, Kevin, et al.
Published: (2025)
by: Tan, Kevin, et al.
Published: (2025)
Sample-Efficient Constrained Reinforcement Learning with General Parameterization
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
LANTERN: LLM-Augmented Neurosymbolic Transfer with Experience-Gated Reasoning Networks
by: Alinejad, Mahyar, et al.
Published: (2026)
by: Alinejad, Mahyar, et al.
Published: (2026)
Oracle-Robust Online Alignment for Large Language Models
by: Li, Zimeng, et al.
Published: (2026)
by: Li, Zimeng, et al.
Published: (2026)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023)
by: Mondal, Washim Uddin, et al.
Published: (2023)
AI Cap-and-Trade: Efficiency Incentives for Accessibility and Sustainability
by: Bornstein, Marco, et al.
Published: (2026)
by: Bornstein, Marco, et al.
Published: (2026)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
A Technical Survey of Reinforcement Learning Techniques for Large Language Models
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
by: Zhang, Jiefu, et al.
Published: (2026)
by: Zhang, Jiefu, et al.
Published: (2026)
A Scalable Quantum Non-local Neural Network for Image Classification
by: Gupta, Sparsh, et al.
Published: (2024)
by: Gupta, Sparsh, et al.
Published: (2024)
A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Network Diffuser for Placing-Scheduling Service Function Chains with Inverse Demonstration
by: Zhang, Zuyuan, et al.
Published: (2025)
by: Zhang, Zuyuan, et al.
Published: (2025)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions
by: Fourati, Fares, et al.
Published: (2025)
by: Fourati, Fares, et al.
Published: (2025)
Stronger Approximation Guarantees for Non-Monotone γ-Weakly DR-Submodular Maximization
by: Jadav, Hareshkumar, et al.
Published: (2026)
by: Jadav, Hareshkumar, et al.
Published: (2026)
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
by: Pedramfar, Mohammad, et al.
Published: (2026)
by: Pedramfar, Mohammad, et al.
Published: (2026)
Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
by: Moradipari, Ahmadreza, et al.
Published: (2023)
by: Moradipari, Ahmadreza, et al.
Published: (2023)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
by: Pedramfar, Mohammad, et al.
Published: (2023)
by: Pedramfar, Mohammad, et al.
Published: (2023)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
by: Reddy, Avinash, et al.
Published: (2026)
by: Reddy, Avinash, et al.
Published: (2026)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
Variational Offline Multi-agent Skill Discovery
by: Chen, Jiayu, et al.
Published: (2024)
by: Chen, Jiayu, et al.
Published: (2024)
A Bi-directional Quantum Search Algorithm
by: Konar, Debanjan, et al.
Published: (2024)
by: Konar, Debanjan, et al.
Published: (2024)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
Similar Items
-
On The Global Convergence Of Online RLHF With Neural Parametrization
by: Gaur, Mudit, et al.
Published: (2024) -
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
by: Gaur, Mudit, et al.
Published: (2025) -
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
by: Gaur, Mudit, et al.
Published: (2025) -
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
by: Gaur, Mudit, et al.
Published: (2025) -
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)