Long-term Safe Reinforcement Learning with Binary Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Wachi, Akifumi, Hashimoto, Wataru, Hashimoto, Kazumune |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
Target Return Optimizer for Multi-Game Decision Transformer
by: Tatematsu, Kensuke, et al.
Published: (2025)
by: Tatematsu, Kensuke, et al.
Published: (2025)
A Survey of Constraint Formulations in Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
Long and Short-Term Constraints Driven Safe Reinforcement Learning for Autonomous Driving
by: Hu, Xuemin, et al.
Published: (2024)
by: Hu, Xuemin, et al.
Published: (2024)
Learning-based Event-triggered MPC with Gaussian processes under terminal constraints
by: Onoue, Yuga, et al.
Published: (2021)
by: Onoue, Yuga, et al.
Published: (2021)
Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning
by: Zhao, Weiye, et al.
Published: (2024)
by: Zhao, Weiye, et al.
Published: (2024)
Sampling-Based Safe Reinforcement Learning
by: Vignola, Luca, et al.
Published: (2026)
by: Vignola, Luca, et al.
Published: (2026)
GUARD: A Safe Reinforcement Learning Benchmark
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Revisiting Safe Exploration in Safe Reinforcement learning
by: Eckel, David, et al.
Published: (2024)
by: Eckel, David, et al.
Published: (2024)
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
by: Su, Huikang, et al.
Published: (2025)
by: Su, Huikang, et al.
Published: (2025)
Leveraging Analytic Gradients in Provably Safe Reinforcement Learning
by: Walter, Tim, et al.
Published: (2025)
by: Walter, Tim, et al.
Published: (2025)
Safe Reinforcement Learning in a Simulated Robotic Arm
by: Kovač, Luka, et al.
Published: (2023)
by: Kovač, Luka, et al.
Published: (2023)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
by: Zheng, Yinan, et al.
Published: (2024)
by: Zheng, Yinan, et al.
Published: (2024)
Safe Offline Reinforcement Learning with Real-Time Budget Constraints
by: Lin, Qian, et al.
Published: (2023)
by: Lin, Qian, et al.
Published: (2023)
Physics-model-guided Worst-case Sampling for Safe Reinforcement Learning
by: Cao, Hongpeng, et al.
Published: (2024)
by: Cao, Hongpeng, et al.
Published: (2024)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
by: Kanai, Sekitoshi, et al.
Published: (2025)
by: Kanai, Sekitoshi, et al.
Published: (2025)
TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
by: Tomov, Momchil S., et al.
Published: (2025)
by: Tomov, Momchil S., et al.
Published: (2025)
Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2026)
by: Brahmanage, Janaka Chathuranga, et al.
Published: (2026)
HAIM-DRL: Enhanced Human-in-the-loop Reinforcement Learning for Safe and Efficient Autonomous Driving
by: Huang, Zilin, et al.
Published: (2024)
by: Huang, Zilin, et al.
Published: (2024)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026)
by: Wachi, Akifumi, et al.
Published: (2026)
Real-World Offline Reinforcement Learning from Vision Language Model Feedback
by: Venkataraman, Sreyas, et al.
Published: (2024)
by: Venkataraman, Sreyas, et al.
Published: (2024)
Trustworthy Human-AI Collaboration: Reinforcement Learning with Human Feedback and Physics Knowledge for Safe Autonomous Driving
by: Huang, Zilin, et al.
Published: (2024)
by: Huang, Zilin, et al.
Published: (2024)
Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
by: Li, Dianzhao, et al.
Published: (2025)
by: Li, Dianzhao, et al.
Published: (2025)
Beyond Scalar Rewards: Distributional Reinforcement Learning with Preordered Objectives for Safe and Reliable Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2026)
by: Abouelazm, Ahmed, et al.
Published: (2026)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)
by: Kudo, Mikoto, et al.
Published: (2026)
A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering
by: Qi, Qihan, et al.
Published: (2024)
by: Qi, Qihan, et al.
Published: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
Simplex-enabled Safe Continual Learning Machine
by: Cao, Hongpeng, et al.
Published: (2024)
by: Cao, Hongpeng, et al.
Published: (2024)
Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning
by: Spoor, Lindsay, et al.
Published: (2025)
by: Spoor, Lindsay, et al.
Published: (2025)
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
by: Geles, Ismail, et al.
Published: (2026)
by: Geles, Ismail, et al.
Published: (2026)
Integrating LTL Constraints into PPO for Safe Reinforcement Learning
by: Zhang, Maifang, et al.
Published: (2026)
by: Zhang, Maifang, et al.
Published: (2026)
Bayesian Deep Learning for Segmentation for Autonomous Safe Planetary Landing
by: Tomita, Kento, et al.
Published: (2021)
by: Tomita, Kento, et al.
Published: (2021)
GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
by: Zhou, Zhehua, et al.
Published: (2024)
by: Zhou, Zhehua, et al.
Published: (2024)
Adaptive Querying for Reward Learning from Human Feedback
by: Anand, Yashwanthi, et al.
Published: (2024)
by: Anand, Yashwanthi, et al.
Published: (2024)
Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
by: Hashimoto, Wataru, et al.
Published: (2024)
by: Hashimoto, Wataru, et al.
Published: (2024)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
by: Hashimoto, Wataru, et al.
Published: (2024)
by: Hashimoto, Wataru, et al.
Published: (2024)
Safe Deep Policy Adaptation
by: Xiao, Wenli, et al.
Published: (2023)
by: Xiao, Wenli, et al.
Published: (2023)
Mining the Long Tail: A Comparative Study of Data-Centric Criticality Metrics for Robust Offline Reinforcement Learning in Autonomous Motion Planning
by: Guillen-Perez, Antonio
Published: (2025)
by: Guillen-Perez, Antonio
Published: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)
by: Hwang, Minjune, et al.
Published: (2026)
Similar Items
-
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025) -
Target Return Optimizer for Multi-Game Decision Transformer
by: Tatematsu, Kensuke, et al.
Published: (2025) -
A Survey of Constraint Formulations in Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2024) -
Long and Short-Term Constraints Driven Safe Reinforcement Learning for Autonomous Driving
by: Hu, Xuemin, et al.
Published: (2024) -
Learning-based Event-triggered MPC with Gaussian processes under terminal constraints
by: Onoue, Yuga, et al.
Published: (2021)