Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
Fuente:
arXiv
Saved in:
| Main Authors: | Mustafin, Arsenii, Sheng, Xinyi, Baumann, Dominik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Value Iteration Convergence in Connected MDPs
by: Mustafin, Arsenii, et al.
Published: (2024)
by: Mustafin, Arsenii, et al.
Published: (2024)
Analysis of Value Iteration Through Absolute Probability Sequences
by: Mustafin, Arsenii, et al.
Published: (2025)
by: Mustafin, Arsenii, et al.
Published: (2025)
Ergodicity in reinforcement learning
by: Baumann, Dominik, et al.
Published: (2026)
by: Baumann, Dominik, et al.
Published: (2026)
MDP Geometry, Normalization and Reward Balancing Solvers
by: Mustafin, Arsenii, et al.
Published: (2024)
by: Mustafin, Arsenii, et al.
Published: (2024)
Closing the gap between SVRG and TD-SVRG with Gradient Splitting
by: Mustafin, Arsenii, et al.
Published: (2022)
by: Mustafin, Arsenii, et al.
Published: (2022)
Geometric Re-Analysis of Classical MDP Solving Algorithms
by: Mustafin, Arsenii, et al.
Published: (2025)
by: Mustafin, Arsenii, et al.
Published: (2025)
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning
by: Sheng, Xinyi, et al.
Published: (2025)
by: Sheng, Xinyi, et al.
Published: (2025)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
by: Grand-Clément, Julien, et al.
Published: (2023)
by: Grand-Clément, Julien, et al.
Published: (2023)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Bridging the Gap Between Average and Discounted TD Learning
by: Tian, Haoxing, et al.
Published: (2026)
by: Tian, Haoxing, et al.
Published: (2026)
Reducing Reward Dependence in RL Through Adaptive Confidence Discounting
by: Satici, Muhammed Yusuf, et al.
Published: (2025)
by: Satici, Muhammed Yusuf, et al.
Published: (2025)
Analyzing and Bridging the Gap between Maximizing Total Reward and Discounted Reward in Deep Reinforcement Learning
by: Yin, Shuyu, et al.
Published: (2024)
by: Yin, Shuyu, et al.
Published: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Bandit Simulation for Average Reward Inference
by: Praharaj, Samya, et al.
Published: (2026)
by: Praharaj, Samya, et al.
Published: (2026)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2025)
by: Ganesh, Swetha, et al.
Published: (2025)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Reinforcement Learning with LTL and $ω$-Regular Objectives via Optimality-Preserving Translation to Average Rewards
by: Le, Xuan-Bach, et al.
Published: (2024)
by: Le, Xuan-Bach, et al.
Published: (2024)
A Unified Analysis for Finite Weight Averaging
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Implicit Updates for Average-Reward Temporal Difference Learning
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Revisiting Weight Averaging for Model Merging
by: Choi, Jiho, et al.
Published: (2024)
by: Choi, Jiho, et al.
Published: (2024)
Safe reinforcement learning in uncertain contexts
by: Baumann, Dominik, et al.
Published: (2024)
by: Baumann, Dominik, et al.
Published: (2024)
Revisiting Bayesian Model Averaging in the Era of Foundation Models
by: Park, Mijung
Published: (2025)
by: Park, Mijung
Published: (2025)
A Unified Linear Speedup Analysis of Federated Averaging and Nesterov FedAvg
by: Qu, Zhaonan, et al.
Published: (2020)
by: Qu, Zhaonan, et al.
Published: (2020)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023)
by: Mondal, Washim Uddin, et al.
Published: (2023)
Transfer Learning in Latent Contextual Bandits with Covariate Shift Through Causal Transportability
by: Deng, Mingwei, et al.
Published: (2025)
by: Deng, Mingwei, et al.
Published: (2025)
Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator
by: Guo, Zhishuai, et al.
Published: (2021)
by: Guo, Zhishuai, et al.
Published: (2021)
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
Natural Policy Gradient for Average Reward Non-Stationary RL
by: Jali, Neharika, et al.
Published: (2025)
by: Jali, Neharika, et al.
Published: (2025)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
by: Roch, Zachary, et al.
Published: (2025)
by: Roch, Zachary, et al.
Published: (2025)
Logarithmic Regret of Exploration in Average Reward Markov Decision Processes
by: Boone, Victor, et al.
Published: (2025)
by: Boone, Victor, et al.
Published: (2025)
A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Value bounds and Convergence Analysis for Averages of LRP attributions
by: Binder, Alexander, et al.
Published: (2025)
by: Binder, Alexander, et al.
Published: (2025)
Iterative Foundation Model Fine-Tuning on Multiple Rewards
by: Ghari, Pouya M., et al.
Published: (2025)
by: Ghari, Pouya M., et al.
Published: (2025)
Optimistic Dual Averaging Unifies Modern Optimizers
by: Pethick, Thomas, et al.
Published: (2026)
by: Pethick, Thomas, et al.
Published: (2026)
Second-Order Mirror Descent: Convergence in Games Beyond Averaging and Discounting
by: Gao, Bolin, et al.
Published: (2021)
by: Gao, Bolin, et al.
Published: (2021)
Similar Items
-
On Value Iteration Convergence in Connected MDPs
by: Mustafin, Arsenii, et al.
Published: (2024) -
Analysis of Value Iteration Through Absolute Probability Sequences
by: Mustafin, Arsenii, et al.
Published: (2025) -
Ergodicity in reinforcement learning
by: Baumann, Dominik, et al.
Published: (2026) -
MDP Geometry, Normalization and Reward Balancing Solvers
by: Mustafin, Arsenii, et al.
Published: (2024) -
Closing the gap between SVRG and TD-SVRG with Gradient Splitting
by: Mustafin, Arsenii, et al.
Published: (2022)