Learning Safety Constraints from Demonstrations with Unknown Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Lindner, David, Chen, Xin, Tschiatschek, Sebastian, Hofmann, Katja, Krause, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Safety Constraints for Large Language Models
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why
by: Li, Chenhao, et al.
Published: (2025)
by: Li, Chenhao, et al.
Published: (2025)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
by: Islam, Md Mirajul, et al.
Published: (2026)
by: Islam, Md Mirajul, et al.
Published: (2026)
Learning Causally Invariant Reward Functions from Diverse Demonstrations
by: Ovinnikov, Ivan, et al.
Published: (2024)
by: Ovinnikov, Ivan, et al.
Published: (2024)
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
by: Rocamonde, Juan, et al.
Published: (2023)
by: Rocamonde, Juan, et al.
Published: (2023)
Plasticity Loss in Deep Reinforcement Learning: A Survey
by: Klein, Timo, et al.
Published: (2024)
by: Klein, Timo, et al.
Published: (2024)
Online Learning with Unknown Constraints
by: Sridharan, Karthik, et al.
Published: (2024)
by: Sridharan, Karthik, et al.
Published: (2024)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
by: Farquhar, Sebastian, et al.
Published: (2025)
by: Farquhar, Sebastian, et al.
Published: (2025)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
by: Karimi, Zohre, et al.
Published: (2024)
by: Karimi, Zohre, et al.
Published: (2024)
Probabilistic Constraint for Safety-Critical Reinforcement Learning
by: Chen, Weiqin, et al.
Published: (2023)
by: Chen, Weiqin, et al.
Published: (2023)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
by: Kim, Kihyun, et al.
Published: (2024)
by: Kim, Kihyun, et al.
Published: (2024)
Learning to Explore with Lagrangians for Bandits under Unknown Linear Constraints
by: Das, Udvas, et al.
Published: (2024)
by: Das, Udvas, et al.
Published: (2024)
Safety Modulation: Enhancing Safety in Reinforcement Learning through Cost-Modulated Rewards
by: Zhang, Hanping, et al.
Published: (2025)
by: Zhang, Hanping, et al.
Published: (2025)
Gram: Assessing sabotage propensities via automated alignment auditing
by: Lindner, David, et al.
Published: (2026)
by: Lindner, David, et al.
Published: (2026)
Positive-Unlabeled Constraint Learning for Inferring Nonlinear Continuous Constraints Functions from Expert Demonstrations
by: Peng, Baiyu, et al.
Published: (2024)
by: Peng, Baiyu, et al.
Published: (2024)
Understanding and Improving Hyperbolic Deep Reinforcement Learning
by: Klein, Timo, et al.
Published: (2025)
by: Klein, Timo, et al.
Published: (2025)
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
by: Kim, Kihyun, et al.
Published: (2026)
by: Kim, Kihyun, et al.
Published: (2026)
Diffusion Models as Constrained Samplers for Optimization with Unknown Constraints
by: Kong, Lingkai, et al.
Published: (2024)
by: Kong, Lingkai, et al.
Published: (2024)
Exploratory Machine Learning with Unknown Unknowns
by: Zhao, Peng, et al.
Published: (2020)
by: Zhao, Peng, et al.
Published: (2020)
Distributed Risk-Sensitive Safety Filters for Uncertain Discrete-Time Systems
by: Lederer, Armin, et al.
Published: (2025)
by: Lederer, Armin, et al.
Published: (2025)
Exploring and Addressing Reward Confusion in Offline Preference Learning
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
Conservative Distributional Reinforcement Learning with Safety Constraints
by: Zhang, Hengrui, et al.
Published: (2022)
by: Zhang, Hengrui, et al.
Published: (2022)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
by: Gumbsch, Christian, et al.
Published: (2026)
by: Gumbsch, Christian, et al.
Published: (2026)
Learning Constraint Network from Demonstrations via Positive-Unlabeled Learning with Memory Replay
by: Peng, Baiyu, et al.
Published: (2024)
by: Peng, Baiyu, et al.
Published: (2024)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
by: Cho, Dongkyu Derek, et al.
Published: (2025)
by: Cho, Dongkyu Derek, et al.
Published: (2025)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
by: Low, Siow Meng, et al.
Published: (2024)
by: Low, Siow Meng, et al.
Published: (2024)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
by: Yang, Daniel, et al.
Published: (2026)
by: Yang, Daniel, et al.
Published: (2026)
Optimizing Earth Observation Satellite Schedules under Unknown Operational Constraints: An Active Constraint Acquisition Approach
by: Belaid, Mohamed-Bachir
Published: (2026)
by: Belaid, Mohamed-Bachir
Published: (2026)
Breaking the Reclustering Barrier in Centroid-based Deep Clustering
by: Miklautz, Lukas, et al.
Published: (2024)
by: Miklautz, Lukas, et al.
Published: (2024)
When a Reinforcement Learning Agent Encounters Unknown Unknowns
by: Zhu, Juntian, et al.
Published: (2025)
by: Zhu, Juntian, et al.
Published: (2025)
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025)
by: Krause, Andreas, et al.
Published: (2025)
An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
by: Plesner, Andreas, et al.
Published: (2026)
by: Plesner, Andreas, et al.
Published: (2026)
Towards the Transferability of Rewards Recovered via Regularized Inverse Reinforcement Learning
by: Schlaginhaufen, Andreas, et al.
Published: (2024)
by: Schlaginhaufen, Andreas, et al.
Published: (2024)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Exploring Open-world Continual Learning with Knowns-Unknowns Knowledge Transfer
by: Li, Yujie, et al.
Published: (2025)
by: Li, Yujie, et al.
Published: (2025)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Similar Items
-
Learning Safety Constraints for Large Language Models
by: Chen, Xin, et al.
Published: (2025) -
Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why
by: Li, Chenhao, et al.
Published: (2025) -
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025) -
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
by: Islam, Md Mirajul, et al.
Published: (2026) -
Learning Causally Invariant Reward Functions from Diverse Demonstrations
by: Ovinnikov, Ivan, et al.
Published: (2024)