Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Stradi, Francesco Emanuele, Castiglioni, Matteo, Marchesi, Alberto, Gatti, Nicola |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Adversarial MDPs with Stochastic Hard Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Data-Dependent Regret Bounds for Constrained MABs
by: Genalti, Gianmarco, et al.
Published: (2025)
by: Genalti, Gianmarco, et al.
Published: (2025)
A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints
by: Germano, Jacopo, et al.
Published: (2023)
by: Germano, Jacopo, et al.
Published: (2023)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
by: Stradi, Francesco Emanuele, et al.
Published: (2025)
by: Stradi, Francesco Emanuele, et al.
Published: (2025)
Truly Adapting to Adversarial Constraints in Constrained MABs
by: Stradi, Francesco Emanuele, et al.
Published: (2026)
by: Stradi, Francesco Emanuele, et al.
Published: (2026)
Toward Optimal Regret in Robust Pricing: Decoupling Corruption and Time
by: Kalupahana, Kalana, et al.
Published: (2026)
by: Kalupahana, Kalana, et al.
Published: (2026)
Learning Constrained Markov Decision Processes With Non-stationary Rewards and Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Markov Persuasion Processes: Learning to Persuade from Scratch
by: Bacchiocchi, Francesco, et al.
Published: (2024)
by: Bacchiocchi, Francesco, et al.
Published: (2024)
Replicable Constrained Bandits
by: Bollini, Matteo, et al.
Published: (2026)
by: Bollini, Matteo, et al.
Published: (2026)
Beyond Slater's Condition in Online CMDPs with Stochastic and Adversarial Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2025)
by: Stradi, Francesco Emanuele, et al.
Published: (2025)
Regret Minimization for Piecewise Linear Rewards: Contracts, Auctions, and Beyond
by: Bacchiocchi, Francesco, et al.
Published: (2025)
by: Bacchiocchi, Francesco, et al.
Published: (2025)
Multi-Armed Bandits With Best-Action Queries
by: Bacchiocchi, Francesco, et al.
Published: (2026)
by: Bacchiocchi, Francesco, et al.
Published: (2026)
Better Regret Rates in Bilateral Trade via Sublinear Budget Violation
by: Lunghi, Anna, et al.
Published: (2025)
by: Lunghi, Anna, et al.
Published: (2025)
Learning Optimal Contracts: How to Exploit Small Action Spaces
by: Bacchiocchi, Francesco, et al.
Published: (2023)
by: Bacchiocchi, Francesco, et al.
Published: (2023)
Regret Minimization in Bilateral Trade With Perturbed Markets
by: Lunghi, Anna, et al.
Published: (2026)
by: Lunghi, Anna, et al.
Published: (2026)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
The Sample Complexity of Uniform Approximation for Multi-Dimensional CDFs and Fixed-Price Mechanisms
by: Castiglioni, Matteo, et al.
Published: (2026)
by: Castiglioni, Matteo, et al.
Published: (2026)
Constrained Phi-Equilibria
by: Bernasconi, Martino, et al.
Published: (2023)
by: Bernasconi, Martino, et al.
Published: (2023)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
by: Bernasconi, Martino, et al.
Published: (2024)
by: Bernasconi, Martino, et al.
Published: (2024)
Safe Online Bid Optimization with Return on Investment and Budget Constraints
by: Castiglioni, Matteo, et al.
Published: (2022)
by: Castiglioni, Matteo, et al.
Published: (2022)
Online Resource Allocation With General Constraints
by: Chiefari, Eleonora Fidelia, et al.
Published: (2026)
by: Chiefari, Eleonora Fidelia, et al.
Published: (2026)
A Primal-Dual Online Learning Approach for Dynamic Pricing of Sequentially Displayed Complementary Items under Sale Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Learning in Bayesian Stackelberg Games With Unknown Follower's Types
by: Bollini, Matteo, et al.
Published: (2026)
by: Bollini, Matteo, et al.
Published: (2026)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Universal Dynamic Regret and Constraint Violation Bounds for Constrained Online Convex Optimization
by: Supantha, Subhamon, et al.
Published: (2025)
by: Supantha, Subhamon, et al.
Published: (2025)
No-Regret Learning in Bilateral Trade via Global Budget Balance
by: Bernasconi, Martino, et al.
Published: (2023)
by: Bernasconi, Martino, et al.
Published: (2023)
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
by: Genalti, Gianmarco, et al.
Published: (2024)
by: Genalti, Gianmarco, et al.
Published: (2024)
The Sample Complexity of Stackelberg Games
by: Bacchiocchi, Francesco, et al.
Published: (2024)
by: Bacchiocchi, Francesco, et al.
Published: (2024)
Contracting With a Reinforcement Learning Agent by Playing Trick or Treat
by: Bollini, Matteo, et al.
Published: (2024)
by: Bollini, Matteo, et al.
Published: (2024)
Online Bayesian Persuasion Without a Clue
by: Bacchiocchi, Francesco, et al.
Published: (2024)
by: Bacchiocchi, Francesco, et al.
Published: (2024)
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025)
by: Sahu, Sharan
Published: (2025)
$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2023)
by: Genalti, Gianmarco, et al.
Published: (2023)
Similar Items
-
Learning Adversarial MDPs with Stochastic Hard Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024) -
Data-Dependent Regret Bounds for Constrained MABs
by: Genalti, Gianmarco, et al.
Published: (2025) -
A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints
by: Germano, Jacopo, et al.
Published: (2023) -
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
by: Stradi, Francesco Emanuele, et al.
Published: (2024) -
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
by: Stradi, Francesco Emanuele, et al.
Published: (2025)