Guardado en:
| Autores principales: | Jiang, Jiashuo, Ye, Yinyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2402.16324 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimal Sample Complexity for Average Reward Markov Decision Processes
por: Wang, Shengbo, et al.
Publicado: (2023)
por: Wang, Shengbo, et al.
Publicado: (2023)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
por: Shen, Xun, et al.
Publicado: (2024)
por: Shen, Xun, et al.
Publicado: (2024)
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
por: Zeng, Sihan, et al.
Publicado: (2021)
por: Zeng, Sihan, et al.
Publicado: (2021)
The Value of Information in Resource-Constrained Pricing
por: Ao, Ruicheng, et al.
Publicado: (2026)
por: Ao, Ruicheng, et al.
Publicado: (2026)
Wait-Less Offline Tuning and Re-solving for Online Decision Making
por: Sun, Jingruo, et al.
Publicado: (2024)
por: Sun, Jingruo, et al.
Publicado: (2024)
Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time
por: Mazumdar, Abhijit, et al.
Publicado: (2024)
por: Mazumdar, Abhijit, et al.
Publicado: (2024)
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
por: Kitamura, Toshinori, et al.
Publicado: (2024)
por: Kitamura, Toshinori, et al.
Publicado: (2024)
Beyond $\mathcal{O}(\sqrt{T})$ Regret: Decoupling Learning and Decision-making in Online Linear Programming
por: Gao, Wenzhi, et al.
Publicado: (2025)
por: Gao, Wenzhi, et al.
Publicado: (2025)
Decoupling Learning and Decision-Making: Breaking the $\mathcal{O}(\sqrt{T})$ Barrier in Online Resource Allocation with First-Order Methods
por: Gao, Wenzhi, et al.
Publicado: (2024)
por: Gao, Wenzhi, et al.
Publicado: (2024)
Weakly Time-Coupled Approximation of Markov Decision Processes
por: Soheili, Negar, et al.
Publicado: (2026)
por: Soheili, Negar, et al.
Publicado: (2026)
Online Markov Decision Processes with Terminal Law Constraints
por: Moreno, Bianca Marin, et al.
Publicado: (2026)
por: Moreno, Bianca Marin, et al.
Publicado: (2026)
Online Linear Programming with Batching
por: Xu, Haoran, et al.
Publicado: (2024)
por: Xu, Haoran, et al.
Publicado: (2024)
A Single-Loop Robust Policy Gradient Method for Robust Markov Decision Processes
por: Lin, Zhenwei, et al.
Publicado: (2024)
por: Lin, Zhenwei, et al.
Publicado: (2024)
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
por: Leon, Vincent, et al.
Publicado: (2023)
por: Leon, Vincent, et al.
Publicado: (2023)
Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets
por: Ho, Chin Pang, et al.
Publicado: (2026)
por: Ho, Chin Pang, et al.
Publicado: (2026)
Non-stationary and Varying-discounting Markov Decision Processes for Reinforcement Learning
por: Chen, Zhizuo, et al.
Publicado: (2025)
por: Chen, Zhizuo, et al.
Publicado: (2025)
Robustifying Conditional Portfolio Decisions via Optimal Transport
por: Nguyen, Viet Anh, et al.
Publicado: (2021)
por: Nguyen, Viet Anh, et al.
Publicado: (2021)
Non-Stationary Online Resource Allocation: Learning from a Single Sample
por: Feng, Yiding, et al.
Publicado: (2026)
por: Feng, Yiding, et al.
Publicado: (2026)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
por: Wan, Yi, et al.
Publicado: (2024)
por: Wan, Yi, et al.
Publicado: (2024)
Risk-sensitive Markov Decision Process and Learning under General Utility Functions
por: Wu, Zhengqi, et al.
Publicado: (2023)
por: Wu, Zhengqi, et al.
Publicado: (2023)
Gradient Methods with Online Scaling
por: Gao, Wenzhi, et al.
Publicado: (2024)
por: Gao, Wenzhi, et al.
Publicado: (2024)
Gradient Methods with Online Scaling Part I. Theoretical Foundations
por: Gao, Wenzhi, et al.
Publicado: (2025)
por: Gao, Wenzhi, et al.
Publicado: (2025)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
por: Chu, Ya-Chi, et al.
Publicado: (2025)
por: Chu, Ya-Chi, et al.
Publicado: (2025)
Gradient Methods with Online Scaling Part II. Practical Aspects
por: Chu, Ya-Chi, et al.
Publicado: (2025)
por: Chu, Ya-Chi, et al.
Publicado: (2025)
Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
por: Sattar, Yahya, et al.
Publicado: (2021)
por: Sattar, Yahya, et al.
Publicado: (2021)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
por: Chen, Zixi, et al.
Publicado: (2025)
por: Chen, Zixi, et al.
Publicado: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
por: Wang, Meixuan, et al.
Publicado: (2025)
por: Wang, Meixuan, et al.
Publicado: (2025)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
por: Wang, Shengbo, et al.
Publicado: (2025)
por: Wang, Shengbo, et al.
Publicado: (2025)
Learning Sequential Decisions from Multiple Sources via Group-Robust Markov Decision Processes
por: Xu, Mingyuan, et al.
Publicado: (2026)
por: Xu, Mingyuan, et al.
Publicado: (2026)
An LP-Based Approach for Bilinear Saddle Point Problem with Instance-dependent Guarantee and Noisy Feedback
por: Jiang, Jiashuo, et al.
Publicado: (2026)
por: Jiang, Jiashuo, et al.
Publicado: (2026)
Bayesian Ambiguity Contraction-based Adaptive Robust Markov Decision Processes for Adversarial Surveillance Missions
por: Choi, Jimin, et al.
Publicado: (2025)
por: Choi, Jimin, et al.
Publicado: (2025)
Learning to Price with Resource Constraints: From Full Information to Machine-Learned Prices
por: Ao, Ruicheng, et al.
Publicado: (2025)
por: Ao, Ruicheng, et al.
Publicado: (2025)
SPABA: A Single-Loop and Probabilistic Stochastic Bilevel Algorithm Achieving Optimal Sample Complexity
por: Chu, Tianshu, et al.
Publicado: (2024)
por: Chu, Tianshu, et al.
Publicado: (2024)
Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
por: Hamza, Ishaq, et al.
Publicado: (2026)
por: Hamza, Ishaq, et al.
Publicado: (2026)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
por: Zhang, Weitong, et al.
Publicado: (2021)
por: Zhang, Weitong, et al.
Publicado: (2021)
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
por: Li, Wenye, et al.
Publicado: (2025)
por: Li, Wenye, et al.
Publicado: (2025)
Data-driven Mixed Integer Optimization through Probabilistic Multi-variable Branching
por: Chen, Yanguang, et al.
Publicado: (2023)
por: Chen, Yanguang, et al.
Publicado: (2023)
A Homogenization Approach for Gradient-Dominated Stochastic Optimization
por: Tan, Jiyuan, et al.
Publicado: (2023)
por: Tan, Jiyuan, et al.
Publicado: (2023)
Fast Nonlinear Two-Time-Scale Stochastic Approximation: Achieving $O(1/k)$ Finite-Sample Complexity
por: Doan, Thinh T.
Publicado: (2024)
por: Doan, Thinh T.
Publicado: (2024)
Quantum Algorithms for Bandits with Knapsacks with Improved Regret and Time Complexities
por: Su, Yuexin, et al.
Publicado: (2025)
por: Su, Yuexin, et al.
Publicado: (2025)
Ejemplares similares
-
Optimal Sample Complexity for Average Reward Markov Decision Processes
por: Wang, Shengbo, et al.
Publicado: (2023) -
Flipping-based Policy for Chance-Constrained Markov Decision Processes
por: Shen, Xun, et al.
Publicado: (2024) -
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
por: Zeng, Sihan, et al.
Publicado: (2021) -
The Value of Information in Resource-Constrained Pricing
por: Ao, Ruicheng, et al.
Publicado: (2026) -
Wait-Less Offline Tuning and Re-solving for Online Decision Making
por: Sun, Jingruo, et al.
Publicado: (2024)