Reinforcement Learning by Guided Safe Exploration
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Qisong, Simão, Thiago D., Jansen, Nils, Tindemans, Simon H., Spaan, Matthijs T. J. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023)
by: Zanger, Moritz A., et al.
Published: (2023)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)
by: Weltevrede, Max, et al.
Published: (2025)
Robust Probabilistic Shielding for Safe Offline Reinforcement Learning
by: Galesloot, Maris F. L., et al.
Published: (2026)
by: Galesloot, Maris F. L., et al.
Published: (2026)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
Positive Experience Reflection for Agents in Interactive Text Environments
by: Lippmann, Philip, et al.
Published: (2024)
by: Lippmann, Philip, et al.
Published: (2024)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
by: van der Vaart, Pascal R., et al.
Published: (2025)
by: van der Vaart, Pascal R., et al.
Published: (2025)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
Maintenance Strategies for Sewer Pipes with Multi-State Degradation and Deep Reinforcement Learning
by: Jimenez-Roa, Lisandro A., et al.
Published: (2024)
by: Jimenez-Roa, Lisandro A., et al.
Published: (2024)
Enhance Exploration in Safe Reinforcement Learning with Contrastive Representation Learning
by: Doan, Duc Kien, et al.
Published: (2025)
by: Doan, Duc Kien, et al.
Published: (2025)
Revisiting Safe Exploration in Safe Reinforcement learning
by: Eckel, David, et al.
Published: (2024)
by: Eckel, David, et al.
Published: (2024)
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
by: Zheng, Yinan, et al.
Published: (2024)
by: Zheng, Yinan, et al.
Published: (2024)
Guiding Exploration in Reinforcement Learning Through LLM-Augmented Observations
by: Jain, Vaibhav, et al.
Published: (2025)
by: Jain, Vaibhav, et al.
Published: (2025)
Epigraph-Guided Flow Matching for Safe and Performant Offline Reinforcement Learning
by: Tayal, Manan, et al.
Published: (2026)
by: Tayal, Manan, et al.
Published: (2026)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
by: Ribeiro, João G., et al.
Published: (2025)
by: Ribeiro, João G., et al.
Published: (2025)
Guided Exploration in Reinforcement Learning via Monte Carlo Critic Optimization
by: Kuznetsov, Igor
Published: (2022)
by: Kuznetsov, Igor
Published: (2022)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
by: Han, Seungyub, et al.
Published: (2026)
by: Han, Seungyub, et al.
Published: (2026)
Safe-Support Q-Learning: Learning without Unsafe Exploration
by: Lim, Yeeun, et al.
Published: (2026)
by: Lim, Yeeun, et al.
Published: (2026)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)
by: Zanger, Moritz A., et al.
Published: (2026)
Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration
by: Le, Tuan, et al.
Published: (2025)
by: Le, Tuan, et al.
Published: (2025)
Sampling-Based Safe Reinforcement Learning
by: Vignola, Luca, et al.
Published: (2026)
by: Vignola, Luca, et al.
Published: (2026)
Probabilistic Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
Guardian: Decoupling Exploration from Safety in Reinforcement Learning
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning
by: Chen, Rufeng, et al.
Published: (2026)
by: Chen, Rufeng, et al.
Published: (2026)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
by: Ji, Jiaming, et al.
Published: (2025)
by: Ji, Jiaming, et al.
Published: (2025)
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
by: Gu, Shangding, et al.
Published: (2022)
by: Gu, Shangding, et al.
Published: (2022)
Universal Value-Function Uncertainties
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Online Optimization for Offline Safe Reinforcement Learning
by: Chemingui, Yassine, et al.
Published: (2025)
by: Chemingui, Yassine, et al.
Published: (2025)
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation
by: Dai, Juntao, et al.
Published: (2024)
by: Dai, Juntao, et al.
Published: (2024)
In-context Exploration-Exploitation for Reinforcement Learning
by: Dai, Zhenwen, et al.
Published: (2024)
by: Dai, Zhenwen, et al.
Published: (2024)
Satisficing Exploration for Deep Reinforcement Learning
by: Arumugam, Dilip, et al.
Published: (2024)
by: Arumugam, Dilip, et al.
Published: (2024)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
by: Tayal, Mumuksh, et al.
Published: (2026)
by: Tayal, Mumuksh, et al.
Published: (2026)
Similar Items
-
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024) -
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024) -
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023) -
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024) -
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)