Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Kaya, Ege C., Ghasemi, Mahsa, Hashemi, Abolfazl |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Localized Distributional Robustness in Submodular Multi-Task Subset Selection
by: Kaya, Ege C., et al.
Published: (2024)
by: Kaya, Ege C., et al.
Published: (2024)
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
by: Fazla, Arda, et al.
Published: (2026)
by: Fazla, Arda, et al.
Published: (2026)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)
by: Zhang, Zhongjun, et al.
Published: (2026)
Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum-Gladyshev Noise
by: Upadhyay, Antesh, et al.
Published: (2026)
by: Upadhyay, Antesh, et al.
Published: (2026)
FedSGM: A Unified Framework for Constraint Aware, Bidirectionally Compressed, Multi-Step Federated Optimization
by: Upadhyay, Antesh, et al.
Published: (2026)
by: Upadhyay, Antesh, et al.
Published: (2026)
Randomized Greedy Methods for Weak Submodular Sensor Selection with Robustness Considerations
by: Kaya, Ege C., et al.
Published: (2024)
by: Kaya, Ege C., et al.
Published: (2024)
Planning and Learning in Average Risk-aware MDPs
by: Wang, Weikai, et al.
Published: (2025)
by: Wang, Weikai, et al.
Published: (2025)
RAMPAGE: RAndomized Mid-Point for debiAsed Gradient Extrapolation
by: Luo, Zhankun, et al.
Published: (2026)
by: Luo, Zhankun, et al.
Published: (2026)
Robust Information Selection for Hypothesis Testing with Misclassification Penalties
by: Bhargav, Jayanth, et al.
Published: (2025)
by: Bhargav, Jayanth, et al.
Published: (2025)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
by: Lee, Kyungbok, et al.
Published: (2026)
by: Lee, Kyungbok, et al.
Published: (2026)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Faster Fixed-Point Methods for Multichain MDPs
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
Submodular Information Selection for Hypothesis Testing with Misclassification Penalties
by: Bhargav, Jayanth, et al.
Published: (2024)
by: Bhargav, Jayanth, et al.
Published: (2024)
Multi-Agent Reinforcement Learning for Joint Police Patrol and Dispatch
by: Repasky, Matthew, et al.
Published: (2024)
by: Repasky, Matthew, et al.
Published: (2024)
Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces
by: Bicer, Osman, et al.
Published: (2025)
by: Bicer, Osman, et al.
Published: (2025)
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023)
by: Mhammedi, Zakaria, et al.
Published: (2023)
Efficient Duple Perturbation Robustness in Low-rank MDPs
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Model approximation in MDPs with unbounded per-step cost
by: Bozkurt, Berk, et al.
Published: (2024)
by: Bozkurt, Berk, et al.
Published: (2024)
Submodular Maximization Approaches for Equitable Client Selection in Federated Learning
by: Jiménez, Andrés Catalino Castillo, et al.
Published: (2024)
by: Jiménez, Andrés Catalino Castillo, et al.
Published: (2024)
Reinforcement Learning in MDPs with Information-Ordered Policies
by: Zhang, Zhongjun, et al.
Published: (2025)
by: Zhang, Zhongjun, et al.
Published: (2025)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Sampling-based Safe Reinforcement Learning for Nonlinear Dynamical Systems
by: Suttle, Wesley A., et al.
Published: (2024)
by: Suttle, Wesley A., et al.
Published: (2024)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2023)
by: Ding, Dongsheng, et al.
Published: (2023)
Sample Efficient Reinforcement Learning with Partial Dynamics Knowledge
by: Alharbi, Meshal, et al.
Published: (2023)
by: Alharbi, Meshal, et al.
Published: (2023)
Deep Reinforcement Learning for Dynamic Order Picking in Warehouse Operations
by: Mahmoudinazlou, Sasan, et al.
Published: (2024)
by: Mahmoudinazlou, Sasan, et al.
Published: (2024)
Reinforcement Learning Approaches for the Orienteering Problem with Stochastic and Dynamic Release Dates
by: Li, Yuanyuan, et al.
Published: (2022)
by: Li, Yuanyuan, et al.
Published: (2022)
Single- vs. Dual-Policy Reinforcement Learning for Dynamic Bike Rebalancing
by: Liang, Jiaqi, et al.
Published: (2024)
by: Liang, Jiaqi, et al.
Published: (2024)
Representative Action Selection for Large Action Space: From Bandits to MDPs
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Generalization Bounds for Sparse Random Feature Expansions
by: Hashemi, Abolfazl, et al.
Published: (2021)
by: Hashemi, Abolfazl, et al.
Published: (2021)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach
by: Qiu, Shuang, et al.
Published: (2022)
by: Qiu, Shuang, et al.
Published: (2022)
Joint Problems in Learning Multiple Dynamical Systems
by: Niu, Mengjia, et al.
Published: (2023)
by: Niu, Mengjia, et al.
Published: (2023)
Joint Learning in the Gaussian Single Index Model
by: Pillaud-Vivien, Loucas, et al.
Published: (2025)
by: Pillaud-Vivien, Loucas, et al.
Published: (2025)
Jointly Computation- and Communication-Efficient Distributed Learning
by: Ren, Xiaoxing, et al.
Published: (2025)
by: Ren, Xiaoxing, et al.
Published: (2025)
Similar Items
-
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
by: Kaya, Ege C., et al.
Published: (2026) -
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
by: Kaya, Ege C., et al.
Published: (2026) -
Localized Distributional Robustness in Submodular Multi-Task Subset Selection
by: Kaya, Ege C., et al.
Published: (2024) -
Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
by: Fazla, Arda, et al.
Published: (2026) -
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)