Safe Learning Under Irreversible Dynamics via Asking for Help
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Plaut, Benjamin, Liévano-Karim, Juan, Zhu, Hanlin, Russell, Stuart |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Avoiding Catastrophe in Online Learning by Asking for Help
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
von: Liaw, Sarah, et al.
Veröffentlicht: (2025)
von: Liaw, Sarah, et al.
Veröffentlicht: (2025)
Getting By Goal Misgeneralization With a Little Help From a Mentor
von: Trinh, Tu, et al.
Veröffentlicht: (2024)
von: Trinh, Tu, et al.
Veröffentlicht: (2024)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
von: Plaut, Benjamin
Veröffentlicht: (2026)
von: Plaut, Benjamin
Veröffentlicht: (2026)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
Learning the Preferences of a Learning Agent
von: Sadek, Karim Abdel, et al.
Veröffentlicht: (2026)
von: Sadek, Karim Abdel, et al.
Veröffentlicht: (2026)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
Cross-Domain Imitation Learning via Optimal Transport
von: Fickinger, Arnaud, et al.
Veröffentlicht: (2021)
von: Fickinger, Arnaud, et al.
Veröffentlicht: (2021)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2023)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2023)
Safe Reinforcement Learning via Recovery-based Shielding with Gaussian Process Dynamics Models
von: Goodall, Alexander W., et al.
Veröffentlicht: (2026)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2026)
Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
Dynamic Model Predictive Shielding for Provably Safe Reinforcement Learning
von: Banerjee, Arko, et al.
Veröffentlicht: (2024)
von: Banerjee, Arko, et al.
Veröffentlicht: (2024)
Verified Safe Reinforcement Learning for Neural Network Dynamic Models
von: Wu, Junlin, et al.
Veröffentlicht: (2024)
von: Wu, Junlin, et al.
Veröffentlicht: (2024)
SMLE: Safe Machine Learning via Embedded Overapproximation
von: Francobaldi, Matteo, et al.
Veröffentlicht: (2024)
von: Francobaldi, Matteo, et al.
Veröffentlicht: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
von: Jenner, Erik, et al.
Veröffentlicht: (2024)
von: Jenner, Erik, et al.
Veröffentlicht: (2024)
Adaptive Spatiotemporal Augmentation for Improving Dynamic Graph Learning
von: Chu, Xu, et al.
Veröffentlicht: (2025)
von: Chu, Xu, et al.
Veröffentlicht: (2025)
Learning to Help in Multi-Class Settings
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
SEAL: Towards Safe Autonomous Driving via Skill-Enabled Adversary Learning for Closed-Loop Scenario Generation
von: Stoler, Benjamin, et al.
Veröffentlicht: (2024)
von: Stoler, Benjamin, et al.
Veröffentlicht: (2024)
Active teacher selection for reward learning
von: Freedman, Rachel, et al.
Veröffentlicht: (2023)
von: Freedman, Rachel, et al.
Veröffentlicht: (2023)
Adaptive Shielding for Safe Reinforcement Learning under Hidden-Parameter Dynamics Shifts
von: Kwon, Minjae, et al.
Veröffentlicht: (2025)
von: Kwon, Minjae, et al.
Veröffentlicht: (2025)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)
von: Lang, Leon, et al.
Veröffentlicht: (2024)
Safe Reinforcement Learning in Black-Box Environments via Adaptive Shielding
von: Bethell, Daniel, et al.
Veröffentlicht: (2024)
von: Bethell, Daniel, et al.
Veröffentlicht: (2024)
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
von: Chen, Keru, et al.
Veröffentlicht: (2024)
von: Chen, Keru, et al.
Veröffentlicht: (2024)
Thermodynamic Irreversibility of Training Algorithms
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
von: Yao, Yihang, et al.
Veröffentlicht: (2023)
von: Yao, Yihang, et al.
Veröffentlicht: (2023)
On Representation Complexity of Model-based and Model-free Reinforcement Learning
von: Zhu, Hanlin, et al.
Veröffentlicht: (2023)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2023)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
von: Wu, David X., et al.
Veröffentlicht: (2025)
von: Wu, David X., et al.
Veröffentlicht: (2025)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
Crafting Interpretable Embeddings by Asking LLMs Questions
von: Benara, Vinamra, et al.
Veröffentlicht: (2024)
von: Benara, Vinamra, et al.
Veröffentlicht: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
von: Jin, Jikai, et al.
Veröffentlicht: (2025)
von: Jin, Jikai, et al.
Veröffentlicht: (2025)
Censoring-Aware Tree-Based Reinforcement Learning for Estimating Dynamic Treatment Regimes with Censored Outcomes
von: Paul, Animesh Kumar, et al.
Veröffentlicht: (2025)
von: Paul, Animesh Kumar, et al.
Veröffentlicht: (2025)
Probabilistic Shielding for Safe Reinforcement Learning
von: Court, Edwin Hamel-De le, et al.
Veröffentlicht: (2025)
von: Court, Edwin Hamel-De le, et al.
Veröffentlicht: (2025)
Reinforcement Learning by Guided Safe Exploration
von: Yang, Qisong, et al.
Veröffentlicht: (2023)
von: Yang, Qisong, et al.
Veröffentlicht: (2023)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025)
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025)
AI Alignment with Changing and Influenceable Reward Functions
von: Carroll, Micah, et al.
Veröffentlicht: (2024)
von: Carroll, Micah, et al.
Veröffentlicht: (2024)
Iterative Batch Reinforcement Learning via Safe Diversified Model-based Policy Search
von: Najib, Amna, et al.
Veröffentlicht: (2024)
von: Najib, Amna, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Avoiding Catastrophe in Online Learning by Asking for Help
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024) -
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
von: Liaw, Sarah, et al.
Veröffentlicht: (2025) -
Getting By Goal Misgeneralization With a Little Help From a Mentor
von: Trinh, Tu, et al.
Veröffentlicht: (2024) -
Safety Training Persists Through Helpfulness Optimization in LLM Agents
von: Plaut, Benjamin
Veröffentlicht: (2026) -
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)