Primal-Dual Policy Optimization for Linear CMDPs with Adversarial Losses
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Kihyun, Bae, Seoungbin, Lee, Dabeen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
Logistic Bandits with $\tilde{O}(\sqrt{dT})$ Regret without Context Diversity Assumptions
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
Neural Logistic Bandits
von: Bae, Seoungbin, et al.
Veröffentlicht: (2025)
von: Bae, Seoungbin, et al.
Veröffentlicht: (2025)
Queue Length Regret Bounds for Contextual Queueing Bandits
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)
Chebyshev Center-Based Direction Selection for Multi-Objective Optimization and Training PINNs
von: Yoon, Hoyeol, et al.
Veröffentlicht: (2026)
von: Yoon, Hoyeol, et al.
Veröffentlicht: (2026)
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
von: Zhu, Jiahui, et al.
Veröffentlicht: (2025)
von: Zhu, Jiahui, et al.
Veröffentlicht: (2025)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
von: Yu, Kihyun, et al.
Veröffentlicht: (2024)
von: Yu, Kihyun, et al.
Veröffentlicht: (2024)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Beyond Slater's Condition in Online CMDPs with Stochastic and Adversarial Constraints
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2025)
Stochastic-Constrained Stochastic Optimization with Markovian Data
von: Kim, Yeongjong, et al.
Veröffentlicht: (2023)
von: Kim, Yeongjong, et al.
Veröffentlicht: (2023)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
Model-Free, Regret-Optimal Best Policy Identification in Online CMDPs
von: Zhou, Zihan, et al.
Veröffentlicht: (2023)
von: Zhou, Zihan, et al.
Veröffentlicht: (2023)
Off-Policy Primal-Dual Safe Reinforcement Learning
von: Wu, Zifan, et al.
Veröffentlicht: (2024)
von: Wu, Zifan, et al.
Veröffentlicht: (2024)
Beyond Primal-Dual Methods in Bandits with Stochastic and Adversarial Constraints
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
Stochastic Smoothed Primal-Dual Algorithms for Nonconvex Optimization with Linear Inequality Constraints
von: Huang, Ruichuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruichuan, et al.
Veröffentlicht: (2025)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Primal-Dual Neural Algorithmic Reasoning
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
von: Ying, Donghao, et al.
Veröffentlicht: (2022)
von: Ying, Donghao, et al.
Veröffentlicht: (2022)
Adversarial Bandit Optimization with Globally Bounded Perturbations to Linear Losses
von: Cheng, Zhuoyu, et al.
Veröffentlicht: (2026)
von: Cheng, Zhuoyu, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Scalable Mean-Field Variational Inference via Preconditioned Primal-Dual Optimization
von: Lyu, Jinhua, et al.
Veröffentlicht: (2026)
von: Lyu, Jinhua, et al.
Veröffentlicht: (2026)
Nearly Optimal Linear Convergence of Stochastic Primal-Dual Methods for Linear Programming
von: Lu, Haihao, et al.
Veröffentlicht: (2021)
von: Lu, Haihao, et al.
Veröffentlicht: (2021)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
Demand Balancing in Primal-Dual Optimization for Blind Network Revenue Management
von: Miao, Sentao, et al.
Veröffentlicht: (2024)
von: Miao, Sentao, et al.
Veröffentlicht: (2024)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
von: Park, Jaehyun, et al.
Veröffentlicht: (2024)
von: Park, Jaehyun, et al.
Veröffentlicht: (2024)
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback
von: Liu, Qiang, et al.
Veröffentlicht: (2026)
von: Liu, Qiang, et al.
Veröffentlicht: (2026)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
Scalable Min-Max Optimization via Primal-Dual Exact Pareto Optimization
von: Park, Sangwoo, et al.
Veröffentlicht: (2025)
von: Park, Sangwoo, et al.
Veröffentlicht: (2025)
Multi-Timescale Primal Dual Hybrid Gradient with Application to Distributed Optimization
von: Zhang, Junhui, et al.
Veröffentlicht: (2025)
von: Zhang, Junhui, et al.
Veröffentlicht: (2025)
Some Primal-Dual Theory for Subgradient Methods for Strongly Convex Optimization
von: Grimmer, Benjamin, et al.
Veröffentlicht: (2023)
von: Grimmer, Benjamin, et al.
Veröffentlicht: (2023)
Continuous Learned Primal Dual
von: Runkel, Christina, et al.
Veröffentlicht: (2024)
von: Runkel, Christina, et al.
Veröffentlicht: (2024)
Adversarial Dual On-Policy Distillation from Expressive Teacher
von: Wan, Zhenglin, et al.
Veröffentlicht: (2026)
von: Wan, Zhenglin, et al.
Veröffentlicht: (2026)
Primal-Dual Methods for Nonsmooth Nonconvex Optimization with Orthogonality Constraints
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2026)
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2026)
Certifying the Right to Be Forgotten: Primal-Dual Optimization for Sample and Label Unlearning in Vertical Federated Learning
von: Jiang, Yu, et al.
Veröffentlicht: (2025)
von: Jiang, Yu, et al.
Veröffentlicht: (2025)
A Primal-Dual-Assisted Penalty Approach to Bilevel Optimization with Coupled Constraints
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2024)
Drago: Primal-Dual Coupled Variance Reduction for Faster Distributionally Robust Optimization
von: Mehta, Ronak, et al.
Veröffentlicht: (2024)
von: Mehta, Ronak, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
von: Yu, Kihyun, et al.
Veröffentlicht: (2026) -
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
von: Yu, Kihyun, et al.
Veröffentlicht: (2026) -
Logistic Bandits with $\tilde{O}(\sqrt{dT})$ Regret without Context Diversity Assumptions
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026) -
Neural Logistic Bandits
von: Bae, Seoungbin, et al.
Veröffentlicht: (2025) -
Queue Length Regret Bounds for Contextual Queueing Bandits
von: Bae, Seoungbin, et al.
Veröffentlicht: (2026)