Guardado en:
| Autores principales: | Tian, Haoxing, Chen, Zaiwei, Paschalidis, Ioannis Ch., Olshevsky, Alex |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.02103 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
One-Shot Averaging for Distributed TD($λ$) Under Markov Sampling
por: Tian, Haoxing, et al.
Publicado: (2024)
por: Tian, Haoxing, et al.
Publicado: (2024)
Closing the gap between SVRG and TD-SVRG with Gradient Splitting
por: Mustafin, Arsenii, et al.
Publicado: (2022)
por: Mustafin, Arsenii, et al.
Publicado: (2022)
On Value Iteration Convergence in Connected MDPs
por: Mustafin, Arsenii, et al.
Publicado: (2024)
por: Mustafin, Arsenii, et al.
Publicado: (2024)
Analysis of Value Iteration Through Absolute Probability Sequences
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
Geometric Re-Analysis of Classical MDP Solving Algorithms
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
MDP Geometry, Normalization and Reward Balancing Solvers
por: Mustafin, Arsenii, et al.
Publicado: (2024)
por: Mustafin, Arsenii, et al.
Publicado: (2024)
Distributionally Robust Learning in Survival Analysis
por: Jin, Yeping, et al.
Publicado: (2025)
por: Jin, Yeping, et al.
Publicado: (2025)
Adversarial Imitation Learning from Visual Observations using Latent Information
por: Giammarino, Vittorio, et al.
Publicado: (2023)
por: Giammarino, Vittorio, et al.
Publicado: (2023)
Visually Robust Adversarial Imitation Learning from Videos with Contrastive Learning
por: Giammarino, Vittorio, et al.
Publicado: (2024)
por: Giammarino, Vittorio, et al.
Publicado: (2024)
Provably Efficient Off-Policy Adversarial Imitation Learning with Convergence Guarantees
por: Chen, Yilei, et al.
Publicado: (2024)
por: Chen, Yilei, et al.
Publicado: (2024)
Multiple-policy Evaluation via Density Estimation
por: Chen, Yilei, et al.
Publicado: (2024)
por: Chen, Yilei, et al.
Publicado: (2024)
Improving Adaptive Online Learning Using Refined Discretization
por: Zhang, Zhiyu, et al.
Publicado: (2023)
por: Zhang, Zhiyu, et al.
Publicado: (2023)
From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes
por: Chen, Zaiwei, et al.
Publicado: (2025)
por: Chen, Zaiwei, et al.
Publicado: (2025)
Distributionally Robust Token Optimization in RLHF
por: Jin, Yeping, et al.
Publicado: (2026)
por: Jin, Yeping, et al.
Publicado: (2026)
DRO-Augment Framework: Robustness by Synergizing Wasserstein Distributionally Robust Optimization and Data Augmentation
por: Hu, Jiaming, et al.
Publicado: (2025)
por: Hu, Jiaming, et al.
Publicado: (2025)
A Model-Based Approach for Improving Reinforcement Learning Efficiency Leveraging Expert Observations
por: Ozcan, Erhan Can, et al.
Publicado: (2024)
por: Ozcan, Erhan Can, et al.
Publicado: (2024)
Network Epidemic Control via Model Predictive Control: Extended Version
por: Talaei, Mahtab, et al.
Publicado: (2026)
por: Talaei, Mahtab, et al.
Publicado: (2026)
Generalized Policy Improvement Algorithms with Theoretically Supported Sample Reuse
por: Queeney, James, et al.
Publicado: (2022)
por: Queeney, James, et al.
Publicado: (2022)
Achieving $\varepsilon^{-2}$ Dependence for Average-Reward Q-Learning with a New Contraction Principle
por: Chen, Zijun, et al.
Publicado: (2026)
por: Chen, Zijun, et al.
Publicado: (2026)
Optimal Transport Perturbations for Safe Reinforcement Learning with Robustness Guarantees
por: Queeney, James, et al.
Publicado: (2023)
por: Queeney, James, et al.
Publicado: (2023)
Convex SGD: Generalization Without Early Stopping
por: Hendrickx, Julien, et al.
Publicado: (2024)
por: Hendrickx, Julien, et al.
Publicado: (2024)
Smooth Ranking SVM via Cutting-Plane Method
por: Ozcan, Erhan Can, et al.
Publicado: (2024)
por: Ozcan, Erhan Can, et al.
Publicado: (2024)
Analyzing and Bridging the Gap between Maximizing Total Reward and Discounted Reward in Deep Reinforcement Learning
por: Yin, Shuyu, et al.
Publicado: (2024)
por: Yin, Shuyu, et al.
Publicado: (2024)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
por: Grand-Clément, Julien, et al.
Publicado: (2023)
por: Grand-Clément, Julien, et al.
Publicado: (2023)
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
por: Nanda, Phalguni, et al.
Publicado: (2025)
por: Nanda, Phalguni, et al.
Publicado: (2025)
Towards General Preference Alignment: Diffusion Models at Nash Equilibrium
por: Hu, Jiaming, et al.
Publicado: (2026)
por: Hu, Jiaming, et al.
Publicado: (2026)
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
por: Moghaddam, Amirreza Neshaei, et al.
Publicado: (2024)
por: Moghaddam, Amirreza Neshaei, et al.
Publicado: (2024)
Network-Based Epidemic Control Through Optimal Travel and Quarantine Management
por: Talaei, Mahtab, et al.
Publicado: (2024)
por: Talaei, Mahtab, et al.
Publicado: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
por: Hamza, Ishaq, et al.
Publicado: (2026)
por: Hamza, Ishaq, et al.
Publicado: (2026)
Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
por: Nanda, Phalguni, et al.
Publicado: (2026)
por: Nanda, Phalguni, et al.
Publicado: (2026)
ASPEST: Bridging the Gap Between Active Learning and Selective Prediction
por: Chen, Jiefeng, et al.
Publicado: (2023)
por: Chen, Jiefeng, et al.
Publicado: (2023)
Bridging the Gap Between Bayesian Deep Learning and Ensemble Weather Forecasts
por: Xiong, Xinlei, et al.
Publicado: (2025)
por: Xiong, Xinlei, et al.
Publicado: (2025)
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization
por: Lei, Xing, et al.
Publicado: (2025)
por: Lei, Xing, et al.
Publicado: (2025)
Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
por: Chen, Zaiwei, et al.
Publicado: (2026)
por: Chen, Zaiwei, et al.
Publicado: (2026)
Sample Complexity of Linear Quadratic Regulator Without Initial Stability
por: Moghaddam, Amirreza Neshaei, et al.
Publicado: (2025)
por: Moghaddam, Amirreza Neshaei, et al.
Publicado: (2025)
Personalized Multi-Agent Average Reward TD-Learning via Joint Linear Approximation
por: Wang, Leo Muxing, et al.
Publicado: (2026)
por: Wang, Leo Muxing, et al.
Publicado: (2026)
Data Deletion Can Help in Adaptive RL
por: Budhraja, Param, et al.
Publicado: (2026)
por: Budhraja, Param, et al.
Publicado: (2026)
Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
Closing the Gap between TD Learning and Supervised Learning -- A Generalisation Point of View
por: Ghugare, Raj, et al.
Publicado: (2024)
por: Ghugare, Raj, et al.
Publicado: (2024)
Ejemplares similares
-
One-Shot Averaging for Distributed TD($λ$) Under Markov Sampling
por: Tian, Haoxing, et al.
Publicado: (2024) -
Closing the gap between SVRG and TD-SVRG with Gradient Splitting
por: Mustafin, Arsenii, et al.
Publicado: (2022) -
On Value Iteration Convergence in Connected MDPs
por: Mustafin, Arsenii, et al.
Publicado: (2024) -
Analysis of Value Iteration Through Absolute Probability Sequences
por: Mustafin, Arsenii, et al.
Publicado: (2025) -
Geometric Re-Analysis of Classical MDP Solving Algorithms
por: Mustafin, Arsenii, et al.
Publicado: (2025)