Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gu, Shangding, Sel, Bilgehan, Ding, Yuhao, Wang, Lu, Lin, Qingwei, Jin, Ming, Knoll, Alois |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
par: Gu, Shangding, et autres
Publié: (2024)
par: Gu, Shangding, et autres
Publié: (2024)
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
par: Gu, Shangding, et autres
Publié: (2024)
par: Gu, Shangding, et autres
Publié: (2024)
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
par: Gu, Shangding, et autres
Publié: (2022)
par: Gu, Shangding, et autres
Publié: (2022)
LLMs Can Plan Only If We Tell Them
par: Sel, Bilgehan, et autres
Publié: (2025)
par: Sel, Bilgehan, et autres
Publié: (2025)
Safe Continual Domain Adaptation after Sim2Real Transfer of Reinforcement Learning Policies in Robotics
par: Josifovski, Josip, et autres
Publié: (2025)
par: Josifovski, Josip, et autres
Publié: (2025)
TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning
par: Gu, Shangding, et autres
Publié: (2024)
par: Gu, Shangding, et autres
Publié: (2024)
A CMDP-within-online framework for Meta-Safe Reinforcement Learning
par: Khattar, Vanshaj, et autres
Publié: (2024)
par: Khattar, Vanshaj, et autres
Publié: (2024)
Reinforcement Learning with Backtracking Feedback
par: Sel, Bilgehan, et autres
Publié: (2026)
par: Sel, Bilgehan, et autres
Publié: (2026)
Backtracking for Safety
par: Sel, Bilgehan, et autres
Publié: (2025)
par: Sel, Bilgehan, et autres
Publié: (2025)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
par: Sel, Bilgehan, et autres
Publié: (2026)
par: Sel, Bilgehan, et autres
Publié: (2026)
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
par: Sel, Bilgehan, et autres
Publié: (2023)
par: Sel, Bilgehan, et autres
Publié: (2023)
Multi-Agent Reinforcement Learning for Autonomous Driving: A Survey
par: Zhang, Ruiqi, et autres
Publié: (2024)
par: Zhang, Ruiqi, et autres
Publié: (2024)
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
par: Gu, Shangding
Publié: (2026)
par: Gu, Shangding
Publié: (2026)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
par: Al-Tawaha, Ahmad, et autres
Publié: (2026)
par: Al-Tawaha, Ahmad, et autres
Publié: (2026)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
par: Sel, Bilgehan, et autres
Publié: (2024)
par: Sel, Bilgehan, et autres
Publié: (2024)
Generating Automotive Code: Large Language Models for Software Development and Verification in Safety-Critical Systems
par: Kirchner, Sven, et autres
Publié: (2025)
par: Kirchner, Sven, et autres
Publié: (2025)
A New Perspective On AI Safety Through Control Theory Methodologies
par: Ullrich, Lars, et autres
Publié: (2025)
par: Ullrich, Lars, et autres
Publié: (2025)
Analysis of Randomization Effects on Sim2Real Transfer in Reinforcement Learning for Robotic Manipulation Tasks
par: Josifovski, Josip, et autres
Publié: (2022)
par: Josifovski, Josip, et autres
Publié: (2022)
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning
par: Gu, Shangding, et autres
Publié: (2025)
par: Gu, Shangding, et autres
Publié: (2025)
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
par: Zheng, Zhi, et autres
Publié: (2024)
par: Zheng, Zhi, et autres
Publié: (2024)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
par: Wang, Yuqing, et autres
Publié: (2025)
par: Wang, Yuqing, et autres
Publié: (2025)
LLM-Empowered Functional Safety and Security by Design in Automotive Systems
par: Petrovic, Nenad, et autres
Publié: (2026)
par: Petrovic, Nenad, et autres
Publié: (2026)
Safety-Aligned 3D Object Detection: Single-Vehicle, Cooperative, and End-to-End Perspectives
par: Liao, Brian Hsuan-Cheng, et autres
Publié: (2026)
par: Liao, Brian Hsuan-Cheng, et autres
Publié: (2026)
State Representations as Incentives for Reinforcement Learning Agents: A Sim2Real Analysis on Robotic Grasping
par: Petropoulakis, Panagiotis, et autres
Publié: (2023)
par: Petropoulakis, Panagiotis, et autres
Publié: (2023)
StyleBench: Evaluating thinking styles in Large Language Models
par: Guo, Junyu, et autres
Publié: (2025)
par: Guo, Junyu, et autres
Publié: (2025)
LLMs Should Express Uncertainty Explicitly
par: Guo, Junyu, et autres
Publié: (2026)
par: Guo, Junyu, et autres
Publié: (2026)
Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization
par: Wang, Yuhao, et autres
Publié: (2025)
par: Wang, Yuhao, et autres
Publié: (2025)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
par: Liu, Xianyang, et autres
Publié: (2026)
par: Liu, Xianyang, et autres
Publié: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
par: Huang, Chenghua, et autres
Publié: (2025)
par: Huang, Chenghua, et autres
Publié: (2025)
Self-Evolved Reward Learning for LLMs
par: Huang, Chenghua, et autres
Publié: (2024)
par: Huang, Chenghua, et autres
Publié: (2024)
Trapezoidal Gradient Descent for Effective Reinforcement Learning in Spiking Networks
par: Pan, Yuhao, et autres
Publié: (2024)
par: Pan, Yuhao, et autres
Publié: (2024)
Position: AI Safety Must Embrace an Antifragile Perspective
par: Jin, Ming, et autres
Publié: (2025)
par: Jin, Ming, et autres
Publié: (2025)
Balancing Rewards in Text Summarization: Multi-Objective Reinforcement Learning via HyperVolume Optimization
par: Song, Junjie, et autres
Publié: (2025)
par: Song, Junjie, et autres
Publié: (2025)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
par: Yao, Yihang, et autres
Publié: (2023)
par: Yao, Yihang, et autres
Publié: (2023)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
par: Low, Siow Meng, et autres
Publié: (2024)
par: Low, Siow Meng, et autres
Publié: (2024)
Reinforcement Learning in a Safety-Embedded MDP with Trajectory Optimization
par: Yang, Fan, et autres
Publié: (2023)
par: Yang, Fan, et autres
Publié: (2023)
Leveraging Analytic Gradients in Provably Safe Reinforcement Learning
par: Walter, Tim, et autres
Publié: (2025)
par: Walter, Tim, et autres
Publié: (2025)
Pretrained Bayesian Non-parametric Knowledge Prior in Robotic Long-Horizon Reinforcement Learning
par: Meng, Yuan, et autres
Publié: (2025)
par: Meng, Yuan, et autres
Publié: (2025)
Safety Modulation: Enhancing Safety in Reinforcement Learning through Cost-Modulated Rewards
par: Zhang, Hanping, et autres
Publié: (2025)
par: Zhang, Hanping, et autres
Publié: (2025)
Safe CoR: A Dual-Expert Approach to Integrating Imitation Learning and Safe Reinforcement Learning Using Constraint Rewards
par: Kwon, Hyeokjin, et autres
Publié: (2024)
par: Kwon, Hyeokjin, et autres
Publié: (2024)
Documents similaires
-
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
par: Gu, Shangding, et autres
Publié: (2024) -
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
par: Gu, Shangding, et autres
Publié: (2024) -
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
par: Gu, Shangding, et autres
Publié: (2022) -
LLMs Can Plan Only If We Tell Them
par: Sel, Bilgehan, et autres
Publié: (2025) -
Safe Continual Domain Adaptation after Sim2Real Transfer of Reinforcement Learning Policies in Robotics
par: Josifovski, Josip, et autres
Publié: (2025)