Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hoang, Huy, Mai, Tien, Varakantham, Pradeep |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2025)
by: Hoang, Huy, et al.
Published: (2025)
Imitating Cost-Constrained Behaviors in Reinforcement Learning
by: Shao, Qian, et al.
Published: (2024)
by: Shao, Qian, et al.
Published: (2024)
Offline Safe Reinforcement Learning Using Trajectory Classification
by: Gong, Ze, et al.
Published: (2024)
by: Gong, Ze, et al.
Published: (2024)
Solving Richly Constrained Reinforcement Learning through State Augmentation and Reward Penalties
by: Jiang, Hao, et al.
Published: (2023)
by: Jiang, Hao, et al.
Published: (2023)
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
by: Ge, Zichang, et al.
Published: (2025)
by: Ge, Zichang, et al.
Published: (2025)
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
by: Lu, Yuxiao, et al.
Published: (2023)
by: Lu, Yuxiao, et al.
Published: (2023)
Regret-Based Defense in Adversarial Reinforcement Learning
by: Belaire, Roman, et al.
Published: (2023)
by: Belaire, Roman, et al.
Published: (2023)
Enhancing the Hierarchical Environment Design via Generative Trajectory Modeling
by: Li, Dexun, et al.
Published: (2023)
by: Li, Dexun, et al.
Published: (2023)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
by: Baharav, Tavor Z., et al.
Published: (2025)
by: Baharav, Tavor Z., et al.
Published: (2025)
Automatic LLM Red Teaming
by: Belaire, Roman, et al.
Published: (2025)
by: Belaire, Roman, et al.
Published: (2025)
On Minimizing Adversarial Counterfactual Error in Adversarial RL
by: Belaire, Roman, et al.
Published: (2024)
by: Belaire, Roman, et al.
Published: (2024)
Vertical Federated Learning in Practice: The Good, the Bad, and the Ugly
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations
by: Bui, The Viet, et al.
Published: (2025)
by: Bui, The Viet, et al.
Published: (2025)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Toward Cost-efficient Adaptive Clinical Trials in Knee Osteoarthritis with Reinforcement Learning
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023)
by: Hu, Hengyuan, et al.
Published: (2023)
On Discovering Algorithms for Adversarial Imitation Learning
by: Chirra, Shashank Reddy, et al.
Published: (2025)
by: Chirra, Shashank Reddy, et al.
Published: (2025)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
Of Good Demons and Bad Angels: Guaranteeing Safe Control under Finite Precision
by: Teuber, Samuel, et al.
Published: (2025)
by: Teuber, Samuel, et al.
Published: (2025)
RILe: Reinforced Imitation Learning
by: Albaba, Mert, et al.
Published: (2024)
by: Albaba, Mert, et al.
Published: (2024)
Do No Harm: A Counterfactual Approach to Safe Reinforcement Learning
by: Vaskov, Sean, et al.
Published: (2024)
by: Vaskov, Sean, et al.
Published: (2024)
Reinforcement Learning via Implicit Imitation Guidance
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Offline Safe Policy Optimization From Heterogeneous Feedback
by: Gong, Ze, et al.
Published: (2025)
by: Gong, Ze, et al.
Published: (2025)
Predicting Bad Goods Risk Scores with ARIMA Time Series: A Novel Risk Assessment Approach
by: Gond, Bishwajit Prasad
Published: (2025)
by: Gond, Bishwajit Prasad
Published: (2025)
Ambient Diffusion Omni: Training Good Models with Bad Data
by: Daras, Giannis, et al.
Published: (2025)
by: Daras, Giannis, et al.
Published: (2025)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
FADE: Why Bad Descriptions Happen to Good Features
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
Adaptive Primal-Dual Method for Safe Reinforcement Learning
by: Chen, Weiqin, et al.
Published: (2024)
by: Chen, Weiqin, et al.
Published: (2024)
Momentum Contrastive Learning with Enhanced Negative Sampling and Hard Negative Filtering
by: Hoang, Duy, et al.
Published: (2025)
by: Hoang, Duy, et al.
Published: (2025)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
Probabilistic Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2026)
by: Burnwal, Returaj, et al.
Published: (2026)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
On Generalization Across Environments In Multi-Objective Reinforcement Learning
by: Teoh, Jayden, et al.
Published: (2025)
by: Teoh, Jayden, et al.
Published: (2025)
A Safe Deep Reinforcement Learning Approach for Energy Efficient Federated Learning in Wireless Communication Networks
by: Koursioumpas, Nikolaos, et al.
Published: (2023)
by: Koursioumpas, Nikolaos, et al.
Published: (2023)
Similar Items
-
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
by: Hoang, Huy, et al.
Published: (2024) -
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024) -
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2025) -
Imitating Cost-Constrained Behaviors in Reinforcement Learning
by: Shao, Qian, et al.
Published: (2024) -
Offline Safe Reinforcement Learning Using Trajectory Classification
by: Gong, Ze, et al.
Published: (2024)