Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tangri, Rohan, Calliess, Jan-Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End Policy Learning of a Statistical Arbitrage Autoencoder Architecture
von: Krause, Fabian, et al.
Veröffentlicht: (2024)
von: Krause, Fabian, et al.
Veröffentlicht: (2024)
Deep Learning for Financial Time Series: A Large-Scale Benchmark of Risk-Adjusted Performance
von: Saly-Kaufmann, Adir, et al.
Veröffentlicht: (2026)
von: Saly-Kaufmann, Adir, et al.
Veröffentlicht: (2026)
Uniform Convergence Beyond Glivenko-Cantelli
von: Devale, Tanmay, et al.
Veröffentlicht: (2025)
von: Devale, Tanmay, et al.
Veröffentlicht: (2025)
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
von: Luis, Carlos E., et al.
Veröffentlicht: (2023)
von: Luis, Carlos E., et al.
Veröffentlicht: (2023)
The Empirical Mean is Minimax Optimal for Local Glivenko-Cantelli
von: Cohen, Doron, et al.
Veröffentlicht: (2024)
von: Cohen, Doron, et al.
Veröffentlicht: (2024)
Glivenko-Cantelli for $f$-divergence
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
PAC-Bayesian Bounds on Constrained f-Entropic Risk Measures
von: Atbir, Hind, et al.
Veröffentlicht: (2025)
von: Atbir, Hind, et al.
Veröffentlicht: (2025)
State-wise Constrained Policy Optimization
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
e-COP : Episodic Constrained Optimization of Policies
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
Proactive Constrained Policy Optimization with Preemptive Penalty
von: Yang, Ning, et al.
Veröffentlicht: (2025)
von: Yang, Ning, et al.
Veröffentlicht: (2025)
Equivariant Goal Conditioned Contrastive Reinforcement Learning
von: Tangri, Arsh, et al.
Veröffentlicht: (2025)
von: Tangri, Arsh, et al.
Veröffentlicht: (2025)
Matrix-Valued Optimism is Matrix-Valued Augmentation: Additive Hybrid Designs for Constrained Optimization
von: Zhao, Jiayi
Veröffentlicht: (2026)
von: Zhao, Jiayi
Veröffentlicht: (2026)
Inference Time Policy Optimization for Offline RL with Differentiable World Models
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
von: Deb, Rohan, et al.
Veröffentlicht: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
Constrained Group Relative Policy Optimization
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
Extreme Value Policy Optimization for Safe Reinforcement Learning
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
Autoregressive Policy Optimization for Constrained Allocation Tasks
von: Winkel, David, et al.
Veröffentlicht: (2024)
von: Winkel, David, et al.
Veröffentlicht: (2024)
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
von: Ghosh, Ayon, et al.
Veröffentlicht: (2024)
von: Ghosh, Ayon, et al.
Veröffentlicht: (2024)
Universal Dynamic Regret and Constraint Violation Bounds for Constrained Online Convex Optimization
von: Supantha, Subhamon, et al.
Veröffentlicht: (2025)
von: Supantha, Subhamon, et al.
Veröffentlicht: (2025)
Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
von: Ying, Chengyang, et al.
Veröffentlicht: (2022)
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
von: Lee, Jane H., et al.
Veröffentlicht: (2025)
von: Lee, Jane H., et al.
Veröffentlicht: (2025)
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
von: Niu, Yifan, et al.
Veröffentlicht: (2025)
von: Niu, Yifan, et al.
Veröffentlicht: (2025)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Massively Scaling Explicit Policy-conditioned Value Functions
von: Bohlinger, Nico, et al.
Veröffentlicht: (2025)
von: Bohlinger, Nico, et al.
Veröffentlicht: (2025)
Equivariant Offline Reinforcement Learning
von: Tangri, Arsh, et al.
Veröffentlicht: (2024)
von: Tangri, Arsh, et al.
Veröffentlicht: (2024)
Co2PO: Coordinated Constrained Policy Optimization for Multi-Agent RL
von: Patel, Shrenik, et al.
Veröffentlicht: (2026)
von: Patel, Shrenik, et al.
Veröffentlicht: (2026)
AlignIQL: Policy Alignment in Implicit Q-Learning through Constrained Optimization
von: He, Longxiang, et al.
Veröffentlicht: (2024)
von: He, Longxiang, et al.
Veröffentlicht: (2024)
Constrained Policy Optimization via Sampling-Based Weight-Space Projection
von: Cao, Shengfan, et al.
Veröffentlicht: (2025)
von: Cao, Shengfan, et al.
Veröffentlicht: (2025)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
Value-Free Policy Optimization via Reward Partitioning
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
von: Si, Wenwen, et al.
Veröffentlicht: (2025)
von: Si, Wenwen, et al.
Veröffentlicht: (2025)
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning
von: Hazra, Somnath, et al.
Veröffentlicht: (2025)
von: Hazra, Somnath, et al.
Veröffentlicht: (2025)
Optimal Bounds for Adversarial Constrained Online Convex Optimization
von: Ferreira, Ricardo N., et al.
Veröffentlicht: (2025)
von: Ferreira, Ricardo N., et al.
Veröffentlicht: (2025)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
Data-Dependent Regret Bounds for Constrained MABs
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)
Stepwise Alignment for Constrained Language Model Policy Optimization
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
Rectified Robust Policy Optimization for Model-Uncertain Constrained Reinforcement Learning without Strong Duality
von: Ma, Shaocong, et al.
Veröffentlicht: (2025)
von: Ma, Shaocong, et al.
Veröffentlicht: (2025)
Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction
von: Zhao, Weiye, et al.
Veröffentlicht: (2024)
von: Zhao, Weiye, et al.
Veröffentlicht: (2024)
Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
von: Wang, Xiaoshuang, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoshuang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
End-to-End Policy Learning of a Statistical Arbitrage Autoencoder Architecture
von: Krause, Fabian, et al.
Veröffentlicht: (2024) -
Deep Learning for Financial Time Series: A Large-Scale Benchmark of Risk-Adjusted Performance
von: Saly-Kaufmann, Adir, et al.
Veröffentlicht: (2026) -
Uniform Convergence Beyond Glivenko-Cantelli
von: Devale, Tanmay, et al.
Veröffentlicht: (2025) -
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
von: Luis, Carlos E., et al.
Veröffentlicht: (2023) -
The Empirical Mean is Minimax Optimal for Local Glivenko-Cantelli
von: Cohen, Doron, et al.
Veröffentlicht: (2024)