Verified Safe Reinforcement Learning for Neural Network Dynamic Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Junlin, Zhang, Huan, Vorobeychik, Yevgeniy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Multi-Agent Reinforcement Learning for Assessing False-Data Injection Attacks on Transportation Networks
by: Eghtesad, Taha, et al.
Published: (2023)
by: Eghtesad, Taha, et al.
Published: (2023)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
by: Lanier, Michael, et al.
Published: (2024)
by: Lanier, Michael, et al.
Published: (2024)
Learning Linear Utility Functions From Pairwise Comparison Queries
by: Ge, Luise, et al.
Published: (2024)
by: Ge, Luise, et al.
Published: (2024)
Online Feedback Efficient Active Target Discovery in Partially Observable Environments
by: Sarkar, Anindya, et al.
Published: (2025)
by: Sarkar, Anindya, et al.
Published: (2025)
Conformal Reachability for Safe Control in Unknown Environments
by: Ma, Xinhang, et al.
Published: (2026)
by: Ma, Xinhang, et al.
Published: (2026)
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees
by: Ma, Xinhang, et al.
Published: (2025)
by: Ma, Xinhang, et al.
Published: (2025)
Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
by: Ge, Luise, et al.
Published: (2025)
by: Ge, Luise, et al.
Published: (2025)
Axioms for AI Alignment from Human Feedback
by: Ge, Luise, et al.
Published: (2024)
by: Ge, Luise, et al.
Published: (2024)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023)
by: Wang, Jiongxiao, et al.
Published: (2023)
Dynamic Model Predictive Shielding for Provably Safe Reinforcement Learning
by: Banerjee, Arko, et al.
Published: (2024)
by: Banerjee, Arko, et al.
Published: (2024)
Adversarial Reinforcement Learning for Detecting False Data Injection Attacks in Vehicular Routing
by: Eghtesad, Taha, et al.
Published: (2026)
by: Eghtesad, Taha, et al.
Published: (2026)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
by: Zhou, Xingjian, et al.
Published: (2024)
by: Zhou, Xingjian, et al.
Published: (2024)
A Scalable Approach to Solving Simulation-Based Network Security Games
by: Lanier, Michael, et al.
Published: (2026)
by: Lanier, Michael, et al.
Published: (2026)
Demonstration Guided Multi-Objective Reinforcement Learning
by: Lu, Junlin, et al.
Published: (2024)
by: Lu, Junlin, et al.
Published: (2024)
Safe Reinforcement Learning via Recovery-based Shielding with Gaussian Process Dynamics Models
by: Goodall, Alexander W., et al.
Published: (2026)
by: Goodall, Alexander W., et al.
Published: (2026)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
by: Cai, Xin-Qiang, et al.
Published: (2025)
by: Cai, Xin-Qiang, et al.
Published: (2025)
Active Geospatial Search for Efficient Tenant Eviction Outreach
by: Sarkar, Anindya, et al.
Published: (2024)
by: Sarkar, Anindya, et al.
Published: (2024)
A Meta-Learning Approach for Multi-Objective Reinforcement Learning in Sustainable Home Environments
by: Lu, Junlin, et al.
Published: (2024)
by: Lu, Junlin, et al.
Published: (2024)
Skill-based Safe Reinforcement Learning with Risk Planning
by: Zhang, Hanping, et al.
Published: (2025)
by: Zhang, Hanping, et al.
Published: (2025)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
by: Cho, Dongkyu Derek, et al.
Published: (2025)
by: Cho, Dongkyu Derek, et al.
Published: (2025)
Adaptive Shielding for Safe Reinforcement Learning under Hidden-Parameter Dynamics Shifts
by: Kwon, Minjae, et al.
Published: (2025)
by: Kwon, Minjae, et al.
Published: (2025)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
Probabilistic Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
by: Lu, Xiaodong, et al.
Published: (2026)
by: Lu, Xiaodong, et al.
Published: (2026)
SafeDreamer: Safe Reinforcement Learning with World Models
by: Huang, Weidong, et al.
Published: (2023)
by: Huang, Weidong, et al.
Published: (2023)
Safe Deep Model-Based Reinforcement Learning with Lyapunov Functions
by: Zhang, Harry
Published: (2024)
by: Zhang, Harry
Published: (2024)
Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial
by: Li, Haoyu, et al.
Published: (2026)
by: Li, Haoyu, et al.
Published: (2026)
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
by: Zhang, Feng, et al.
Published: (2026)
by: Zhang, Feng, et al.
Published: (2026)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
by: Ji, Jiaming, et al.
Published: (2025)
by: Ji, Jiaming, et al.
Published: (2025)
DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
by: Li, Long, et al.
Published: (2026)
by: Li, Long, et al.
Published: (2026)
Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning
by: Chen, Rufeng, et al.
Published: (2026)
by: Chen, Rufeng, et al.
Published: (2026)
Online Optimization for Offline Safe Reinforcement Learning
by: Chemingui, Yassine, et al.
Published: (2025)
by: Chemingui, Yassine, et al.
Published: (2025)
Revisiting Safe Exploration in Safe Reinforcement learning
by: Eckel, David, et al.
Published: (2024)
by: Eckel, David, et al.
Published: (2024)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
by: Tayal, Mumuksh, et al.
Published: (2026)
by: Tayal, Mumuksh, et al.
Published: (2026)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2023)
by: Yao, Yihang, et al.
Published: (2023)
Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning
by: Zhao, Weiye, et al.
Published: (2024)
by: Zhao, Weiye, et al.
Published: (2024)
Similar Items
-
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024) -
Multi-Agent Reinforcement Learning for Assessing False-Data Injection Attacks on Transportation Networks
by: Eghtesad, Taha, et al.
Published: (2023) -
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
by: Lanier, Michael, et al.
Published: (2024) -
Learning Linear Utility Functions From Pairwise Comparison Queries
by: Ge, Luise, et al.
Published: (2024) -
Online Feedback Efficient Active Target Discovery in Partially Observable Environments
by: Sarkar, Anindya, et al.
Published: (2025)