Optimistic critics can empower small actors
Fuente:
arXiv
Saved in:
| Main Authors: | Mastikhina, Olya, Sreenivas, Dhruv, Castro, Pablo Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Lifetime tuning is incompatible with continual reinforcement learning
by: Mesbahi, Golnaz, et al.
Published: (2024)
by: Mesbahi, Golnaz, et al.
Published: (2024)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
by: Moulin, Antoine, et al.
Published: (2025)
by: Moulin, Antoine, et al.
Published: (2025)
Towards Interpretability in Audio and Visual Affective Machine Learning: A Review
by: Johnson, David S., et al.
Published: (2023)
by: Johnson, David S., et al.
Published: (2023)
Optimistic Information Directed Sampling
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
Adversarial Imitation Learning via Boosting
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Optimistic Policy Regularization
by: Pham, Mai, et al.
Published: (2026)
by: Pham, Mai, et al.
Published: (2026)
Sparse Optimistic Information Directed Sampling
by: Schwartz, Ludovic, et al.
Published: (2025)
by: Schwartz, Ludovic, et al.
Published: (2025)
Bayesian policy gradient and actor-critic algorithms
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
SOMBRL: Scalable and Optimistic Model-Based RL
by: Sukhija, Bhavya, et al.
Published: (2025)
by: Sukhija, Bhavya, et al.
Published: (2025)
Optimistic Task Inference for Behavior Foundation Models
by: Rupf, Thomas, et al.
Published: (2025)
by: Rupf, Thomas, et al.
Published: (2025)
Bayesian Optimistic Optimisation with Exponentially Decaying Regret
by: Tran-The, Hung, et al.
Published: (2021)
by: Tran-The, Hung, et al.
Published: (2021)
Optimistic Dual Averaging Unifies Modern Optimizers
by: Pethick, Thomas, et al.
Published: (2026)
by: Pethick, Thomas, et al.
Published: (2026)
Optimistic Learning for Communication Networks
by: Iosifidis, George, et al.
Published: (2025)
by: Iosifidis, George, et al.
Published: (2025)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Optimistic Reinforcement Learning with Quantile Objectives
by: Alipour-Vaezi, Mohammad, et al.
Published: (2025)
by: Alipour-Vaezi, Mohammad, et al.
Published: (2025)
Optimistic Multi-Agent Policy Gradient
by: Zhao, Wenshuai, et al.
Published: (2023)
by: Zhao, Wenshuai, et al.
Published: (2023)
The Formalism-Implementation Gap in Reinforcement Learning Research
by: Castro, Pablo Samuel
Published: (2025)
by: Castro, Pablo Samuel
Published: (2025)
On Stability in Optimistic Bilevel Optimization
by: Royset, Johannes O.
Published: (2024)
by: Royset, Johannes O.
Published: (2024)
Optimistic Interior Point Methods for Sequential Hypothesis Testing by Betting
by: Chen, Can, et al.
Published: (2025)
by: Chen, Can, et al.
Published: (2025)
Optimistic Q-learning for average reward and episodic reinforcement learning
by: Agrawal, Priyank, et al.
Published: (2024)
by: Agrawal, Priyank, et al.
Published: (2024)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Optimistic Algorithms for Adaptive Estimation of the Average Treatment Effect
by: Neopane, Ojash, et al.
Published: (2025)
by: Neopane, Ojash, et al.
Published: (2025)
Tail Distribution of Regret in Optimistic Reinforcement Learning
by: Khodadadian, Sajad, et al.
Published: (2025)
by: Khodadadian, Sajad, et al.
Published: (2025)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
Finite-time analysis of single-timescale actor-critic
by: Chen, Xuyang, et al.
Published: (2022)
by: Chen, Xuyang, et al.
Published: (2022)
A Theory of Optimistically Universal Online Learnability for General Concept Classes
by: Hanneke, Steve, et al.
Published: (2025)
by: Hanneke, Steve, et al.
Published: (2025)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
by: Viel, Stefano, et al.
Published: (2025)
by: Viel, Stefano, et al.
Published: (2025)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning
by: Li, Yao-Hui, et al.
Published: (2026)
by: Li, Yao-Hui, et al.
Published: (2026)
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality
by: Yan, Yu-Hu, et al.
Published: (2025)
by: Yan, Yu-Hu, et al.
Published: (2025)
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
by: Zhu, Jiahui, et al.
Published: (2025)
by: Zhu, Jiahui, et al.
Published: (2025)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
by: Nguyen, Anh Duc, et al.
Published: (2025)
by: Nguyen, Anh Duc, et al.
Published: (2025)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
Optimistic Online Non-stochastic Control via FTRL
by: Mhaisen, Naram, et al.
Published: (2024)
by: Mhaisen, Naram, et al.
Published: (2024)
opp/ai: Optimistic Privacy-Preserving AI on Blockchain
by: So, Cathie, et al.
Published: (2024)
by: So, Cathie, et al.
Published: (2024)
An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints
by: Lekeufack, Jordan, et al.
Published: (2024)
by: Lekeufack, Jordan, et al.
Published: (2024)
Generalizing soft actor-critic algorithms to discrete action spaces
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
Similar Items
-
Position: Lifetime tuning is incompatible with continual reinforcement learning
by: Mesbahi, Golnaz, et al.
Published: (2024) -
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
by: Moulin, Antoine, et al.
Published: (2025) -
Towards Interpretability in Audio and Visual Affective Machine Learning: A Review
by: Johnson, David S., et al.
Published: (2023) -
Optimistic Information Directed Sampling
by: Neu, Gergely, et al.
Published: (2024) -
Adversarial Imitation Learning via Boosting
by: Chang, Jonathan D., et al.
Published: (2024)