Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Navdeep, Dahan, Tehila, Cohen, Lior, Barua, Ananyabrata, Ramponi, Giorgia, Levy, Kfir Yehuda, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025)
von: Koren, Uri, et al.
Veröffentlicht: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
von: Dahan, Tehila, et al.
Veröffentlicht: (2025)
von: Dahan, Tehila, et al.
Veröffentlicht: (2025)
Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum
von: Dahan, Tehila, et al.
Veröffentlicht: (2026)
von: Dahan, Tehila, et al.
Veröffentlicht: (2026)
Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous Training
von: Dahan, Tehila, et al.
Veröffentlicht: (2024)
von: Dahan, Tehila, et al.
Veröffentlicht: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
von: Dahan, Tehila, et al.
Veröffentlicht: (2023)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Local MixVR: Breaking the Communication-Sample Dependence in Distributed Learning
von: Dahan, Tehila, et al.
Veröffentlicht: (2026)
von: Dahan, Tehila, et al.
Veröffentlicht: (2026)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
von: Cohen, Lior, et al.
Veröffentlicht: (2025)
von: Cohen, Lior, et al.
Veröffentlicht: (2025)
Improving Token-Based World Models with Parallel Observation Prediction
von: Cohen, Lior, et al.
Veröffentlicht: (2024)
von: Cohen, Lior, et al.
Veröffentlicht: (2024)
Privacy-Preserving Federated Convex Optimization: Balancing Partial-Participation and Efficiency via Noise Cancellation
von: Reshef, Roie, et al.
Veröffentlicht: (2025)
von: Reshef, Roie, et al.
Veröffentlicht: (2025)
MinMaxMin $Q$-learning
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Representative Action Selection for Large Action Space: From Bandits to MDPs
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
von: Perets, Binyamin, et al.
Veröffentlicht: (2026)
von: Perets, Binyamin, et al.
Veröffentlicht: (2026)
Safe Primal-Dual Optimization with a Single Smooth Constraint
von: Usmanova, Ilnura, et al.
Veröffentlicht: (2025)
von: Usmanova, Ilnura, et al.
Veröffentlicht: (2025)
Efficient Fairness-Performance Pareto Front Computation
von: Kozdoba, Mark, et al.
Veröffentlicht: (2024)
von: Kozdoba, Mark, et al.
Veröffentlicht: (2024)
The Value of Mechanistic Priors in Sequential Decision Making
von: Shufaro, Itai, et al.
Veröffentlicht: (2026)
von: Shufaro, Itai, et al.
Veröffentlicht: (2026)
Semiparametric Robust Estimation of Population Location
von: Barua, Ananyabrata, et al.
Veröffentlicht: (2025)
von: Barua, Ananyabrata, et al.
Veröffentlicht: (2025)
Representation-Driven Reinforcement Learning
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
Sobolev Space Regularised Pre Density Models
von: Kozdoba, Mark, et al.
Veröffentlicht: (2023)
von: Kozdoba, Mark, et al.
Veröffentlicht: (2023)
Representative Action Selection for Large Action Space Bandit Families
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2024)
von: Freihaut, Till, et al.
Veröffentlicht: (2024)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022)
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022)
Spectral Bellman Method: Unifying Representation and Exploration in RL
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
A Classification View on Meta Learning Bandits
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
Gradient-Variation Online Adaptivity for Accelerated Optimization with Hölder Smoothness
von: Zhao, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yuheng, et al.
Veröffentlicht: (2025)
Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models
von: Vainshtein, Ron, et al.
Veröffentlicht: (2025)
von: Vainshtein, Ron, et al.
Veröffentlicht: (2025)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
von: Valensi, David, et al.
Veröffentlicht: (2024)
von: Valensi, David, et al.
Veröffentlicht: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Actor-Critics Can Achieve Optimal Sample Efficiency
von: Tan, Kevin, et al.
Veröffentlicht: (2025)
von: Tan, Kevin, et al.
Veröffentlicht: (2025)
Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning
von: Zhou, Tianchen, et al.
Veröffentlicht: (2024)
von: Zhou, Tianchen, et al.
Veröffentlicht: (2024)
SQT -- std $Q$-target
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
von: Hamza, Ishaq, et al.
Veröffentlicht: (2026)
von: Hamza, Ishaq, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024) -
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025) -
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025) -
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
von: Dahan, Tehila, et al.
Veröffentlicht: (2023) -
Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
von: Dahan, Tehila, et al.
Veröffentlicht: (2025)