Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
Fuente:
arXiv
Saved in:
| Main Authors: | Zuo, Qian, He, Fengxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLM Safety Be Ensured by Constraining Parameter Regions?
by: Li, Zongmin, et al.
Published: (2026)
by: Li, Zongmin, et al.
Published: (2026)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Near-Constant Strong Violation and Last-Iterate Convergence for Online CMDPs via Decaying Safety Margins
by: Zuo, Qian, et al.
Published: (2026)
by: Zuo, Qian, et al.
Published: (2026)
PRISM: Parallel Reward Integration with Symmetry for MORL
by: van der Knaap, Finn, et al.
Published: (2026)
by: van der Knaap, Finn, et al.
Published: (2026)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
by: Tang, Kenton, et al.
Published: (2026)
by: Tang, Kenton, et al.
Published: (2026)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
by: Vora, Manav, et al.
Published: (2024)
by: Vora, Manav, et al.
Published: (2024)
Integrating LTL Constraints into PPO for Safe Reinforcement Learning
by: Zhang, Maifang, et al.
Published: (2026)
by: Zhang, Maifang, et al.
Published: (2026)
XAI for In-hospital Mortality Prediction via Multimodal ICU Data
by: Li, Xingqiao, et al.
Published: (2023)
by: Li, Xingqiao, et al.
Published: (2023)
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)
by: Schnitzer, Yannik, et al.
Published: (2024)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
When a Reinforcement Learning Agent Encounters Unknown Unknowns
by: Zhu, Juntian, et al.
Published: (2025)
by: Zhu, Juntian, et al.
Published: (2025)
Uncertain Multi-Objective Recommendation via Orthogonal Meta-Learning Enhanced Bayesian Optimization
by: Wang, Hongxu, et al.
Published: (2025)
by: Wang, Hongxu, et al.
Published: (2025)
Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
by: Zheng, Haoyang, et al.
Published: (2024)
by: Zheng, Haoyang, et al.
Published: (2024)
Stochastic Variance-Reduced Iterative Hard Thresholding in Graph Sparsity Optimization
by: Fox, Derek, et al.
Published: (2024)
by: Fox, Derek, et al.
Published: (2024)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
by: Zeng, Qirun, et al.
Published: (2025)
by: Zeng, Qirun, et al.
Published: (2025)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks
by: Shu, Aijie, et al.
Published: (2026)
by: Shu, Aijie, et al.
Published: (2026)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
Distributed Risk-Sensitive Safety Filters for Uncertain Discrete-Time Systems
by: Lederer, Armin, et al.
Published: (2025)
by: Lederer, Armin, et al.
Published: (2025)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Geometry of Drifting MDPs with Path-Integral Stability Certificates
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
Stochastic Penalty-Barrier Methods for Constrained Machine Learning
by: Bosák, Adam, et al.
Published: (2026)
by: Bosák, Adam, et al.
Published: (2026)
Formal Verification of Graph Convolutional Networks with Uncertain Node Features and Uncertain Graph Structure
by: Ladner, Tobias, et al.
Published: (2024)
by: Ladner, Tobias, et al.
Published: (2024)
Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
by: Ge, Luise, et al.
Published: (2025)
by: Ge, Luise, et al.
Published: (2025)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Risk-averse Total-reward MDPs with ERM and EVaR
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026)
by: Shah, Anvay, et al.
Published: (2026)
On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
by: Dong, Zixuan, et al.
Published: (2022)
by: Dong, Zixuan, et al.
Published: (2022)
Drawback of Enforcing Equivariance and its Compensation via the Lens of Expressive Power
by: Chen, Yuzhu, et al.
Published: (2025)
by: Chen, Yuzhu, et al.
Published: (2025)
Ensured: Explanations for Decreasing the Epistemic Uncertainty in Predictions
by: Löfström, Helena, et al.
Published: (2024)
by: Löfström, Helena, et al.
Published: (2024)
STORI: A Benchmark and Taxonomy for Stochastic Environments
by: Barsainyan, Aryan Amit, et al.
Published: (2025)
by: Barsainyan, Aryan Amit, et al.
Published: (2025)
Similar Items
-
Can LLM Safety Be Ensured by Constraining Parameter Regions?
by: Li, Zongmin, et al.
Published: (2026) -
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
by: Cao, Chentao, et al.
Published: (2025) -
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024) -
Near-Constant Strong Violation and Last-Iterate Convergence for Online CMDPs via Decaying Safety Margins
by: Zuo, Qian, et al.
Published: (2026) -
PRISM: Parallel Reward Integration with Symmetry for MORL
by: van der Knaap, Finn, et al.
Published: (2026)