On the Convergence of Single-Timescale Actor-Critic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Navdeep, Agrawal, Priyank, Ramponi, Giorgia, Levy, Kfir Yehuda, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025)
von: Koren, Uri, et al.
Veröffentlicht: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
Privacy-Preserving Federated Convex Optimization: Balancing Partial-Participation and Efficiency via Noise Cancellation
von: Reshef, Roie, et al.
Veröffentlicht: (2025)
von: Reshef, Roie, et al.
Veröffentlicht: (2025)
MinMaxMin $Q$-learning
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Representative Action Selection for Large Action Space: From Bandits to MDPs
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
von: Perets, Binyamin, et al.
Veröffentlicht: (2026)
von: Perets, Binyamin, et al.
Veröffentlicht: (2026)
Efficient Fairness-Performance Pareto Front Computation
von: Kozdoba, Mark, et al.
Veröffentlicht: (2024)
von: Kozdoba, Mark, et al.
Veröffentlicht: (2024)
The Value of Mechanistic Priors in Sequential Decision Making
von: Shufaro, Itai, et al.
Veröffentlicht: (2026)
von: Shufaro, Itai, et al.
Veröffentlicht: (2026)
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2024)
von: Freihaut, Till, et al.
Veröffentlicht: (2024)
Representation-Driven Reinforcement Learning
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
Sobolev Space Regularised Pre Density Models
von: Kozdoba, Mark, et al.
Veröffentlicht: (2023)
von: Kozdoba, Mark, et al.
Veröffentlicht: (2023)
Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
Representative Action Selection for Large Action Space Bandit Families
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms
von: Panda, Prashansa, et al.
Veröffentlicht: (2023)
von: Panda, Prashansa, et al.
Veröffentlicht: (2023)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
Spectral Bellman Method: Unifying Representation and Exploration in RL
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
A Classification View on Meta Learning Bandits
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
von: Mutti, Mirco, et al.
Veröffentlicht: (2025)
Optimistic Q-learning for average reward and episodic reinforcement learning
von: Agrawal, Priyank, et al.
Veröffentlicht: (2024)
von: Agrawal, Priyank, et al.
Veröffentlicht: (2024)
Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation
von: Panda, Prashansa, et al.
Veröffentlicht: (2024)
von: Panda, Prashansa, et al.
Veröffentlicht: (2024)
Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models
von: Vainshtein, Ron, et al.
Veröffentlicht: (2025)
von: Vainshtein, Ron, et al.
Veröffentlicht: (2025)
Gradient-Variation Online Adaptivity for Accelerated Optimization with Hölder Smoothness
von: Zhao, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yuheng, et al.
Veröffentlicht: (2025)
Improving Token-Based World Models with Parallel Observation Prediction
von: Cohen, Lior, et al.
Veröffentlicht: (2024)
von: Cohen, Lior, et al.
Veröffentlicht: (2024)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
von: Valensi, David, et al.
Veröffentlicht: (2024)
von: Valensi, David, et al.
Veröffentlicht: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Segment Feedback
von: Du, Yihan, et al.
Veröffentlicht: (2025)
von: Du, Yihan, et al.
Veröffentlicht: (2025)
SQT -- std $Q$-target
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
von: Soffair, Nitsan, et al.
Veröffentlicht: (2024)
Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator
von: Chen, Xuyang, et al.
Veröffentlicht: (2025)
von: Chen, Xuyang, et al.
Veröffentlicht: (2025)
Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
von: Macuglia, Maël, et al.
Veröffentlicht: (2025)
von: Macuglia, Maël, et al.
Veröffentlicht: (2025)
Q-learning with Posterior Sampling
von: Agrawal, Priyank, et al.
Veröffentlicht: (2025)
von: Agrawal, Priyank, et al.
Veröffentlicht: (2025)
Policy Gradient with Tree Expansion
von: Dalal, Gal, et al.
Veröffentlicht: (2023)
von: Dalal, Gal, et al.
Veröffentlicht: (2023)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
von: Cohen, Lior, et al.
Veröffentlicht: (2025)
von: Cohen, Lior, et al.
Veröffentlicht: (2025)
A Single-Loop Deep Actor-Critic Algorithm for Constrained Reinforcement Learning with Provable Convergence
von: Wang, Kexuan, et al.
Veröffentlicht: (2023)
von: Wang, Kexuan, et al.
Veröffentlicht: (2023)
Safe Primal-Dual Optimization with a Single Smooth Constraint
von: Usmanova, Ilnura, et al.
Veröffentlicht: (2025)
von: Usmanova, Ilnura, et al.
Veröffentlicht: (2025)
Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous Training
von: Dahan, Tehila, et al.
Veröffentlicht: (2024)
von: Dahan, Tehila, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026) -
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025) -
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025) -
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023) -
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)