Policy Gradient Algorithms in Average-Reward Multichain MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jongmin, Ryu, Ernest K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision Processes
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Faster Fixed-Point Methods for Multichain MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
Query-Efficient Zeroth-Order Algorithms for Nonconvex Constrained Optimization
von: Jin, Ruiyang, et al.
Veröffentlicht: (2025)
von: Jin, Ruiyang, et al.
Veröffentlicht: (2025)
Convergence Analyses of Davis-Yin Splitting via Scaled Relative Graphs
von: Lee, Jongmin, et al.
Veröffentlicht: (2022)
von: Lee, Jongmin, et al.
Veröffentlicht: (2022)
Optimal Sensor and Actuator Selection for Factored Markov Decision Processes: Complexity, Approximability and Algorithms
von: Bhargav, Jayanth, et al.
Veröffentlicht: (2024)
von: Bhargav, Jayanth, et al.
Veröffentlicht: (2024)
Constrained Average-Reward Intermittently Observable MDPs
von: Avrachenkov, Konstantin, et al.
Veröffentlicht: (2025)
von: Avrachenkov, Konstantin, et al.
Veröffentlicht: (2025)
Hardness of some optimization problems over correlation polyhedra
von: Caprara, Alberto, et al.
Veröffentlicht: (2026)
von: Caprara, Alberto, et al.
Veröffentlicht: (2026)
On the Induced Norms of Matrices and Grothendieck problems
von: Truong, Lan V., et al.
Veröffentlicht: (2026)
von: Truong, Lan V., et al.
Veröffentlicht: (2026)
Constrained Nonnegative Gram Feasibility is $\exists\mathbb{R}$-Complete
von: Majumdar, Angshul
Veröffentlicht: (2026)
von: Majumdar, Angshul
Veröffentlicht: (2026)
On Big-M Reformulations of Bilevel Linear Programs: Hardness of A Posteriori Verification
von: Ketkov, Sergey S., et al.
Veröffentlicht: (2026)
von: Ketkov, Sergey S., et al.
Veröffentlicht: (2026)
Information Redistribution Under Reductions in NP Search
von: Wei, Jing-Yuan
Veröffentlicht: (2026)
von: Wei, Jing-Yuan
Veröffentlicht: (2026)
On the Complexity of p-Order Cone Programs
von: Blanco, Víctor, et al.
Veröffentlicht: (2025)
von: Blanco, Víctor, et al.
Veröffentlicht: (2025)
On a class of interdiction problems with partition matroids: complexity and polynomial-time algorithms
von: Ketkov, Sergey S., et al.
Veröffentlicht: (2024)
von: Ketkov, Sergey S., et al.
Veröffentlicht: (2024)
Efficient LP warmstarting for linear modifications of the constraint matrix
von: Derval, Guillaume, et al.
Veröffentlicht: (2025)
von: Derval, Guillaume, et al.
Veröffentlicht: (2025)
Avoiding Deadlocks via Weak Deadlock Sets
von: Oriolo, Gianpaolo, et al.
Veröffentlicht: (2024)
von: Oriolo, Gianpaolo, et al.
Veröffentlicht: (2024)
On the Degree Automatability of Sum-of-Squares Proofs
von: Bortolotti, Alex, et al.
Veröffentlicht: (2025)
von: Bortolotti, Alex, et al.
Veröffentlicht: (2025)
Parameterized complexity of scheduling unit-time jobs with generalized precedence constraints
von: Büsing, Christina, et al.
Veröffentlicht: (2025)
von: Büsing, Christina, et al.
Veröffentlicht: (2025)
A System-Dynamic Based Simulation and Bayesian Optimization for Inventory Management
von: Maitra, Sarit
Veröffentlicht: (2024)
von: Maitra, Sarit
Veröffentlicht: (2024)
Learning complexity of gradient descent and conjugate gradient algorithms
von: Jiao, Xianqi, et al.
Veröffentlicht: (2024)
von: Jiao, Xianqi, et al.
Veröffentlicht: (2024)
Benchmarking of Quantum and Classical Computing in Large-Scale Dynamic Portfolio Optimization Under Market Frictions
von: Chen, Ying, et al.
Veröffentlicht: (2025)
von: Chen, Ying, et al.
Veröffentlicht: (2025)
Geometric and computational hardness of bilevel programming
von: Bolte, Jérôme, et al.
Veröffentlicht: (2024)
von: Bolte, Jérôme, et al.
Veröffentlicht: (2024)
Counterfactual Explanations for Integer Optimization Problems
von: Engelhardt, Felix, et al.
Veröffentlicht: (2025)
von: Engelhardt, Felix, et al.
Veröffentlicht: (2025)
A parameterized linear formulation of the integer hull
von: Eisenbrand, Friedrich, et al.
Veröffentlicht: (2025)
von: Eisenbrand, Friedrich, et al.
Veröffentlicht: (2025)
The Complexity of Computing KKT Solutions of Quadratic Programs
von: Fearnley, John, et al.
Veröffentlicht: (2023)
von: Fearnley, John, et al.
Veröffentlicht: (2023)
Reduction from the partition problem: Dynamic lot sizing problem with polynomial complexity
von: Sim, Chee-Khian
Veröffentlicht: (2024)
von: Sim, Chee-Khian
Veröffentlicht: (2024)
Tight Time Complexities in Parallel Stochastic Optimization with Arbitrary Computation Dynamics
von: Tyurin, Alexander
Veröffentlicht: (2024)
von: Tyurin, Alexander
Veröffentlicht: (2024)
Iterative Optimization of Multidimensional Functions on Turing Machines under Performance Guarantees
von: Boche, Holger, et al.
Veröffentlicht: (2025)
von: Boche, Holger, et al.
Veröffentlicht: (2025)
A parallel framework for graphical optimal transport
von: Fan, Jiaojiao, et al.
Veröffentlicht: (2024)
von: Fan, Jiaojiao, et al.
Veröffentlicht: (2024)
The Complexity of Recognizing Facets for the Knapsack Polytope
von: Chen, Rui, et al.
Veröffentlicht: (2022)
von: Chen, Rui, et al.
Veröffentlicht: (2022)
Centrality of shortest paths: Algorithms and complexity results
von: Phosavanh, Johnson, et al.
Veröffentlicht: (2024)
von: Phosavanh, Johnson, et al.
Veröffentlicht: (2024)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
Deterministic Algorithm for Non-monotone Submodular Maximization under Matroid and Knapsack Constraints
von: Chen, Shengminjie, et al.
Veröffentlicht: (2026)
von: Chen, Shengminjie, et al.
Veröffentlicht: (2026)
Real Stability and Log Concavity are coNP-Hard
von: Chin, Tracy
Veröffentlicht: (2024)
von: Chin, Tracy
Veröffentlicht: (2024)
Two Choices are Enough for P-LCPs, USOs, and Colorful Tangents
von: Borzechowski, Michaela, et al.
Veröffentlicht: (2024)
von: Borzechowski, Michaela, et al.
Veröffentlicht: (2024)
Point Convergence of Nesterov's Accelerated Gradient Method: An AI-Assisted Proof
von: Jang, Uijeong, et al.
Veröffentlicht: (2025)
von: Jang, Uijeong, et al.
Veröffentlicht: (2025)
Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Algorithms for Standard-form ILP Problems via Komlós' Discrepancy Setting
von: Gribanov, Dmitry, et al.
Veröffentlicht: (2026)
von: Gribanov, Dmitry, et al.
Veröffentlicht: (2026)
Deflated Dynamics Value Iteration
von: Lee, Jongmin, et al.
Veröffentlicht: (2024)
von: Lee, Jongmin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision Processes
von: Lee, Jongmin, et al.
Veröffentlicht: (2025) -
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
von: Lee, Jongmin, et al.
Veröffentlicht: (2025) -
Faster Fixed-Point Methods for Multichain MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2025) -
Query-Efficient Zeroth-Order Algorithms for Nonconvex Constrained Optimization
von: Jin, Ruiyang, et al.
Veröffentlicht: (2025) -
Convergence Analyses of Davis-Yin Splitting via Scaled Relative Graphs
von: Lee, Jongmin, et al.
Veröffentlicht: (2022)