Smoothed functional-based gradient algorithms for off-policy reinforcement learning: A non-asymptotic viewpoint
Fuente:
arXiv
Saved in:
| Main Authors: | Vijayan, Nithia, A, Prashanth L. |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A policy gradient approach for optimization of smooth risk measures
by: Vijayan, Nithia, et al.
Published: (2022)
by: Vijayan, Nithia, et al.
Published: (2022)
Policy Gradient Methods for Distortion Risk Measures
by: Vijayan, Nithia, et al.
Published: (2021)
by: Vijayan, Nithia, et al.
Published: (2021)
Optimization of utility-based shortfall risk: A non-asymptotic viewpoint
by: Gupte, Sumedh, et al.
Published: (2023)
by: Gupte, Sumedh, et al.
Published: (2023)
Self-Interested Agents in Collaborative Machine Learning: An Incentivized Adaptive Data-Centric Framework
by: Vijayan, Nithia, et al.
Published: (2024)
by: Vijayan, Nithia, et al.
Published: (2024)
An advantage based policy transfer algorithm for reinforcement learning with measures of transferability
by: Alam, Md Ferdous, et al.
Published: (2023)
by: Alam, Md Ferdous, et al.
Published: (2023)
Counterfactual experience augmented off-policy reinforcement learning
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Catastrophic-risk-aware reinforcement learning with extreme-value-theory-based policy gradients
by: Davar, Parisa, et al.
Published: (2024)
by: Davar, Parisa, et al.
Published: (2024)
Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
by: Guan, Yang, et al.
Published: (2021)
by: Guan, Yang, et al.
Published: (2021)
Towards minimax optimal algorithms for Active Simple Hypothesis Testing
by: Vijayan, Sushant
Published: (2025)
by: Vijayan, Sushant
Published: (2025)
Bayesian policy gradient and actor-critic algorithms
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
by: Ghavamzadeh, Mohammad, et al.
Published: (2026)
Independent policy gradient-based reinforcement learning for economic and reliable energy management of multi-microgrid systems
by: Hu, Junkai, et al.
Published: (2025)
by: Hu, Junkai, et al.
Published: (2025)
Risk-sensitive reinforcement learning using expectiles, shortfall risk and optimized certainty equivalent risk
by: Gupte, Sumedh, et al.
Published: (2026)
by: Gupte, Sumedh, et al.
Published: (2026)
Automating proton PBS treatment planning for head and neck cancers using policy gradient-based deep reinforcement learning
by: Wang, Qingqing, et al.
Published: (2024)
by: Wang, Qingqing, et al.
Published: (2024)
Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching
by: Denkert, Robert, et al.
Published: (2024)
by: Denkert, Robert, et al.
Published: (2024)
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
by: Liu, Jiashun, et al.
Published: (2025)
by: Liu, Jiashun, et al.
Published: (2025)
Bellman operator convergence enhancements in reinforcement learning algorithms
by: Kadurha, David Krame, et al.
Published: (2025)
by: Kadurha, David Krame, et al.
Published: (2025)
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
by: Mahajan, Pranav, et al.
Published: (2026)
by: Mahajan, Pranav, et al.
Published: (2026)
Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation
by: Patil, Gandharv, et al.
Published: (2022)
by: Patil, Gandharv, et al.
Published: (2022)
RL-finetuning LLMs from on- and off-policy data with a single algorithm
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Randomized algorithms and PAC bounds for inverse reinforcement learning in continuous spaces
by: Kamoutsi, Angeliki, et al.
Published: (2024)
by: Kamoutsi, Angeliki, et al.
Published: (2024)
How to craft a deep reinforcement learning policy for wind farm flow control
by: Kadoche, Elie, et al.
Published: (2025)
by: Kadoche, Elie, et al.
Published: (2025)
Risk-averse policies for natural gas futures trading using distributional reinforcement learning
by: Hêche, Félicien, et al.
Published: (2025)
by: Hêche, Félicien, et al.
Published: (2025)
Convergence of a model-free entropy-regularized inverse reinforcement learning algorithm
by: Renard, Titouan, et al.
Published: (2024)
by: Renard, Titouan, et al.
Published: (2024)
Full error analysis of policy gradient learning algorithms for exploratory linear quadratic mean-field control problem in continuous time with common noise
by: Frikha, Noufel, et al.
Published: (2024)
by: Frikha, Noufel, et al.
Published: (2024)
Ergodicity in reinforcement learning
by: Baumann, Dominik, et al.
Published: (2026)
by: Baumann, Dominik, et al.
Published: (2026)
A projection-based framework for gradient-free and parallel learning
by: Bergmeister, Andreas, et al.
Published: (2025)
by: Bergmeister, Andreas, et al.
Published: (2025)
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Reward function compression facilitates goal-dependent reinforcement learning
by: Molinaro, Gaia, et al.
Published: (2025)
by: Molinaro, Gaia, et al.
Published: (2025)
Equivalence of stochastic and deterministic policy gradients
by: Todorov, Emo
Published: (2025)
by: Todorov, Emo
Published: (2025)
Policy gradient methods for ordinal policies
by: Weinberger, Simón, et al.
Published: (2025)
by: Weinberger, Simón, et al.
Published: (2025)
BAPR: Bayesian amnesic piecewise-robust reinforcement learning for non-stationary continuous control
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
A first realization of reinforcement learning-based closed-loop EEG-TMS
by: Humaidan, Dania, et al.
Published: (2026)
by: Humaidan, Dania, et al.
Published: (2026)
Self-test loss functions for learning weak-form operators and gradient flows
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Convergence of off-policy TD(0) with linear function approximation for reversible Markov chains
by: Overmars, Maik, et al.
Published: (2025)
by: Overmars, Maik, et al.
Published: (2025)
Q-MARL: A quantum-inspired algorithm using neural message passing for large-scale multi-agent reinforcement learning
by: Vo, Kha, et al.
Published: (2025)
by: Vo, Kha, et al.
Published: (2025)
Transfer learning strategies for accelerating reinforcement-learning-based flow control
by: Salehi, Saeed
Published: (2025)
by: Salehi, Saeed
Published: (2025)
Optimizing Shortfall Risk Metric for Learning Regression Models
by: Ramaswamy, Harish G., et al.
Published: (2025)
by: Ramaswamy, Harish G., et al.
Published: (2025)
Model predictive control-based value estimation for efficient reinforcement learning
by: Wu, Qizhen, et al.
Published: (2023)
by: Wu, Qizhen, et al.
Published: (2023)
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning
by: Sheng, Xinyi, et al.
Published: (2025)
by: Sheng, Xinyi, et al.
Published: (2025)
Similar Items
-
A policy gradient approach for optimization of smooth risk measures
by: Vijayan, Nithia, et al.
Published: (2022) -
Policy Gradient Methods for Distortion Risk Measures
by: Vijayan, Nithia, et al.
Published: (2021) -
Optimization of utility-based shortfall risk: A non-asymptotic viewpoint
by: Gupte, Sumedh, et al.
Published: (2023) -
Self-Interested Agents in Collaborative Machine Learning: An Incentivized Adaptive Data-Centric Framework
by: Vijayan, Nithia, et al.
Published: (2024) -
An advantage based policy transfer algorithm for reinforcement learning with measures of transferability
by: Alam, Md Ferdous, et al.
Published: (2023)