A CMDP-within-online framework for Meta-Safe Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khattar, Vanshaj, Ding, Yuhao, Sel, Bilgehan, Lavaei, Javad, Jin, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimization Solution Functions as Deterministic Policies for Offline Reinforcement Learning
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024)
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024)
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
von: Ying, Donghao, et al.
Veröffentlicht: (2022)
von: Ying, Donghao, et al.
Veröffentlicht: (2022)
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
von: Sel, Bilgehan, et al.
Veröffentlicht: (2023)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2023)
Pausing Policy Learning in Non-stationary Reinforcement Learning
von: Lee, Hyunin, et al.
Veröffentlicht: (2024)
von: Lee, Hyunin, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Backtracking Feedback
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
von: Ding, Yuhao, et al.
Veröffentlicht: (2021)
von: Ding, Yuhao, et al.
Veröffentlicht: (2021)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2026)
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2026)
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Detecting Zero-Day Attacks in Digital Substations via In-Context Learning
von: Manzoor, Faizan, et al.
Veröffentlicht: (2025)
von: Manzoor, Faizan, et al.
Veröffentlicht: (2025)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
von: Sel, Bilgehan, et al.
Veröffentlicht: (2024)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2024)
Provably Efficient Sample Complexity for Robust CMDP
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
StyleBench: Evaluating thinking styles in Large Language Models
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
LLMs Should Express Uncertainty Explicitly
von: Guo, Junyu, et al.
Veröffentlicht: (2026)
von: Guo, Junyu, et al.
Veröffentlicht: (2026)
TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
von: Guo, Junyu, et al.
Veröffentlicht: (2025)
LLMs Can Plan Only If We Tell Them
von: Sel, Bilgehan, et al.
Veröffentlicht: (2025)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2025)
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
von: Gu, Shangding, et al.
Veröffentlicht: (2025)
von: Gu, Shangding, et al.
Veröffentlicht: (2025)
High Probability Complexity Bounds of Trust-Region Stochastic Sequential Quadratic Programming with Heavy-Tailed Noise
von: Fang, Yuchen, et al.
Veröffentlicht: (2025)
von: Fang, Yuchen, et al.
Veröffentlicht: (2025)
Feasibility Consistent Representation Learning for Safe Reinforcement Learning
von: Cen, Zhepeng, et al.
Veröffentlicht: (2024)
von: Cen, Zhepeng, et al.
Veröffentlicht: (2024)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
von: Ma, Ziye, et al.
Veröffentlicht: (2024)
von: Ma, Ziye, et al.
Veröffentlicht: (2024)
Subgradient Method for System Identification with Non-Smooth Objectives
von: Yalcin, Baturalp, et al.
Veröffentlicht: (2025)
von: Yalcin, Baturalp, et al.
Veröffentlicht: (2025)
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
von: Yao, Yihang, et al.
Veröffentlicht: (2024)
von: Yao, Yihang, et al.
Veröffentlicht: (2024)
Controlling Underestimation Bias in Constrained Reinforcement Learning for Safe Exploration
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
Exact Recovery for System Identification with More Corrupt Data than Clean Data
von: Yalcin, Baturalp, et al.
Veröffentlicht: (2023)
von: Yalcin, Baturalp, et al.
Veröffentlicht: (2023)
Extreme Value Policy Optimization for Safe Reinforcement Learning
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
Safe In-Context Reinforcement Learning
von: Moeini, Amir, et al.
Veröffentlicht: (2025)
von: Moeini, Amir, et al.
Veröffentlicht: (2025)
Counterfactually Safe Reinforcement Learning
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
A Trust-Region Interior-Point Stochastic Sequential Quadratic Programming Method
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
A Scalable Approach for Safe and Robust Learning via Lipschitz-Constrained Networks
von: Abdeen, Zain ul, et al.
Veröffentlicht: (2025)
von: Abdeen, Zain ul, et al.
Veröffentlicht: (2025)
Online Bayesian Risk-Averse Reinforcement Learning
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Policy Bifurcation in Safe Reinforcement Learning
von: Zou, Wenjun, et al.
Veröffentlicht: (2024)
von: Zou, Wenjun, et al.
Veröffentlicht: (2024)
Structural Correspondence and Universal Approximation in Diagonal plus Low-Rank Neural Networks
von: Chen, Ying, et al.
Veröffentlicht: (2026)
von: Chen, Ying, et al.
Veröffentlicht: (2026)
Intersection of Reinforcement Learning and Bayesian Optimization for Intelligent Control of Industrial Processes: A Safe MPC-based DPG using Multi-Objective BO
von: Esfahani, Hossein Nejatbakhsh, et al.
Veröffentlicht: (2025)
von: Esfahani, Hossein Nejatbakhsh, et al.
Veröffentlicht: (2025)
Safe RuleFit: Learning Optimal Sparse Rule Model by Meta Safe Screening
von: Kato, Hiroki, et al.
Veröffentlicht: (2018)
von: Kato, Hiroki, et al.
Veröffentlicht: (2018)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
von: Yao, Yihang, et al.
Veröffentlicht: (2023)
von: Yao, Yihang, et al.
Veröffentlicht: (2023)
Conditional Sequence Modeling for Safe Reinforcement Learning
von: Bai, Wensong, et al.
Veröffentlicht: (2026)
von: Bai, Wensong, et al.
Veröffentlicht: (2026)
Vulnerability Analysis of Safe Reinforcement Learning via Inverse Constrained Reinforcement Learning
von: Fan, Jialiang, et al.
Veröffentlicht: (2026)
von: Fan, Jialiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Optimization Solution Functions as Deterministic Policies for Offline Reinforcement Learning
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024) -
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
von: Gu, Shangding, et al.
Veröffentlicht: (2024) -
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024) -
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
von: Ying, Donghao, et al.
Veröffentlicht: (2022) -
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
von: Sel, Bilgehan, et al.
Veröffentlicht: (2023)