Learning Adversarial MDPs with Stochastic Hard Constraints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Stradi, Francesco Emanuele, Castiglioni, Matteo, Marchesi, Alberto, Gatti, Nicola |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Beyond Slater's Condition in Online CMDPs with Stochastic and Adversarial Constraints
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)
Truly Adapting to Adversarial Constraints in Constrained MABs
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2026)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2026)
A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints
von: Germano, Jacopo, et al.
Veröffentlicht: (2023)
von: Germano, Jacopo, et al.
Veröffentlicht: (2023)
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)
Learning Constrained Markov Decision Processes With Non-stationary Rewards and Constraints
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Markov Persuasion Processes: Learning to Persuade from Scratch
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2024)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2024)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Data-Dependent Regret Bounds for Constrained MABs
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)
Multi-Armed Bandits With Best-Action Queries
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2026)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2026)
Toward Optimal Regret in Robust Pricing: Decoupling Corruption and Time
von: Kalupahana, Kalana, et al.
Veröffentlicht: (2026)
von: Kalupahana, Kalana, et al.
Veröffentlicht: (2026)
Replicable Constrained Bandits
von: Bollini, Matteo, et al.
Veröffentlicht: (2026)
von: Bollini, Matteo, et al.
Veröffentlicht: (2026)
Learning Optimal Contracts: How to Exploit Small Action Spaces
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2023)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2023)
Regret Minimization for Piecewise Linear Rewards: Contracts, Auctions, and Beyond
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2025)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2025)
Online Resource Allocation With General Constraints
von: Chiefari, Eleonora Fidelia, et al.
Veröffentlicht: (2026)
von: Chiefari, Eleonora Fidelia, et al.
Veröffentlicht: (2026)
A Primal-Dual Online Learning Approach for Dynamic Pricing of Sequentially Displayed Complementary Items under Sale Constraints
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Beyond Primal-Dual Methods in Bandits with Stochastic and Adversarial Constraints
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
Safe Online Bid Optimization with Return on Investment and Budget Constraints
von: Castiglioni, Matteo, et al.
Veröffentlicht: (2022)
von: Castiglioni, Matteo, et al.
Veröffentlicht: (2022)
The Sample Complexity of Uniform Approximation for Multi-Dimensional CDFs and Fixed-Price Mechanisms
von: Castiglioni, Matteo, et al.
Veröffentlicht: (2026)
von: Castiglioni, Matteo, et al.
Veröffentlicht: (2026)
Learning in Bayesian Stackelberg Games With Unknown Follower's Types
von: Bollini, Matteo, et al.
Veröffentlicht: (2026)
von: Bollini, Matteo, et al.
Veröffentlicht: (2026)
Better Regret Rates in Bilateral Trade via Sublinear Budget Violation
von: Lunghi, Anna, et al.
Veröffentlicht: (2025)
von: Lunghi, Anna, et al.
Veröffentlicht: (2025)
Regret Minimization in Bilateral Trade With Perturbed Markets
von: Lunghi, Anna, et al.
Veröffentlicht: (2026)
von: Lunghi, Anna, et al.
Veröffentlicht: (2026)
Contracting With a Reinforcement Learning Agent by Playing Trick or Treat
von: Bollini, Matteo, et al.
Veröffentlicht: (2024)
von: Bollini, Matteo, et al.
Veröffentlicht: (2024)
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2024)
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2024)
The Sample Complexity of Stackelberg Games
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2024)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2024)
Online Bayesian Persuasion Without a Clue
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2024)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2024)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
Constrained Phi-Equilibria
von: Bernasconi, Martino, et al.
Veröffentlicht: (2023)
von: Bernasconi, Martino, et al.
Veröffentlicht: (2023)
Contract Design Under Approximate Best Responses
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2025)
von: Bacchiocchi, Francesco, et al.
Veröffentlicht: (2025)
Online Learning under Budget and ROI Constraints via Weak Adaptivity
von: Castiglioni, Matteo, et al.
Veröffentlicht: (2023)
von: Castiglioni, Matteo, et al.
Veröffentlicht: (2023)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2024)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2024)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
von: Ito, Shinji, et al.
Veröffentlicht: (2025)
von: Ito, Shinji, et al.
Veröffentlicht: (2025)
No-Regret Reinforcement Learning in Smooth MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
von: Tsuchiya, Taira, et al.
Veröffentlicht: (2025)
von: Tsuchiya, Taira, et al.
Veröffentlicht: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
Projection by Convolution: Optimal Sample Complexity for Reinforcement Learning in Continuous-Space MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Eluder-based Regret for Stochastic Contextual MDPs
von: Levy, Orin, et al.
Veröffentlicht: (2022)
von: Levy, Orin, et al.
Veröffentlicht: (2022)
Probabilistic Iterative Hard Thresholding for Sparse Learning
von: Bergamaschi, Matteo, et al.
Veröffentlicht: (2024)
von: Bergamaschi, Matteo, et al.
Veröffentlicht: (2024)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024) -
Beyond Slater's Condition in Online CMDPs with Stochastic and Adversarial Constraints
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025) -
Truly Adapting to Adversarial Constraints in Constrained MABs
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2026) -
A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints
von: Germano, Jacopo, et al.
Veröffentlicht: (2023) -
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)