Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Di, Qiwei, Zhao, Heyang, He, Jiafan, Gu, Quanquan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
di: Zhao, Heyang, et al.
Pubblicazione: (2023)
di: Zhao, Heyang, et al.
Pubblicazione: (2023)
Reinforcement Learning from Human Feedback with Active Queries
di: Ji, Kaixuan, et al.
Pubblicazione: (2024)
di: Ji, Kaixuan, et al.
Pubblicazione: (2024)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
di: Zhang, Weitong, et al.
Pubblicazione: (2021)
di: Zhang, Weitong, et al.
Pubblicazione: (2021)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
di: Di, Qiwei, et al.
Pubblicazione: (2023)
di: Di, Qiwei, et al.
Pubblicazione: (2023)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
di: Li, Xuheng, et al.
Pubblicazione: (2024)
di: Li, Xuheng, et al.
Pubblicazione: (2024)
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
di: Li, Runjia, et al.
Pubblicazione: (2024)
di: Li, Runjia, et al.
Pubblicazione: (2024)
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
di: Zhang, Shiyuan, et al.
Pubblicazione: (2026)
di: Zhang, Shiyuan, et al.
Pubblicazione: (2026)
Outlier-robust Autocovariance Least Square Estimation via Iteratively Reweighted Least Square
di: Li, Jiahong, et al.
Pubblicazione: (2026)
di: Li, Jiahong, et al.
Pubblicazione: (2026)
Global Convergence of Iteratively Reweighted Least Squares for Robust Subspace Recovery
di: Lerman, Gilad, et al.
Pubblicazione: (2025)
di: Lerman, Gilad, et al.
Pubblicazione: (2025)
Iterative Pre-Conditioning for Expediting the Gradient-Descent Method: The Distributed Linear Least-Squares Problem
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2020)
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2020)
Recovering Simultaneously Structured Data via Non-Convex Iteratively Reweighted Least Squares
di: Kümmerle, Christian, et al.
Pubblicazione: (2023)
di: Kümmerle, Christian, et al.
Pubblicazione: (2023)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
di: Li, Xuheng, et al.
Pubblicazione: (2025)
di: Li, Xuheng, et al.
Pubblicazione: (2025)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
di: Li, Xuheng, et al.
Pubblicazione: (2025)
di: Li, Xuheng, et al.
Pubblicazione: (2025)
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
di: MacDonald, Lachlan Ewen, et al.
Pubblicazione: (2025)
di: MacDonald, Lachlan Ewen, et al.
Pubblicazione: (2025)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
di: Di, Qiwei, et al.
Pubblicazione: (2024)
di: Di, Qiwei, et al.
Pubblicazione: (2024)
Residuals-based Offline Reinforcement Learning
di: Zhu, Qing, et al.
Pubblicazione: (2026)
di: Zhu, Qing, et al.
Pubblicazione: (2026)
Online Nonstochastic Prediction: Logarithmic Regret via Predictive Online Least Squares
di: Pai, Chih-Fan, et al.
Pubblicazione: (2026)
di: Pai, Chih-Fan, et al.
Pubblicazione: (2026)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
di: Zhang, Junkai, et al.
Pubblicazione: (2023)
di: Zhang, Junkai, et al.
Pubblicazione: (2023)
MARS-M: When Variance Reduction Meets Matrices
di: Liu, Yifeng, et al.
Pubblicazione: (2025)
di: Liu, Yifeng, et al.
Pubblicazione: (2025)
Offline Reinforcement Learning via Inverse Optimization
di: Dimanidis, Ioannis, et al.
Pubblicazione: (2025)
di: Dimanidis, Ioannis, et al.
Pubblicazione: (2025)
Robust Least-Squares Optimization for Data-Driven Predictive Control: A Geometric Approach
di: Bharadwaj, Shreyas, et al.
Pubblicazione: (2025)
di: Bharadwaj, Shreyas, et al.
Pubblicazione: (2025)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
di: Zhang, Zhongjun, et al.
Pubblicazione: (2026)
di: Zhang, Zhongjun, et al.
Pubblicazione: (2026)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
di: Zhou, Angela
Pubblicazione: (2024)
di: Zhou, Angela
Pubblicazione: (2024)
On Regularization via Early Stopping for Least Squares Regression
di: Sonthalia, Rishi, et al.
Pubblicazione: (2024)
di: Sonthalia, Rishi, et al.
Pubblicazione: (2024)
Offline Hierarchical Reinforcement Learning via Inverse Optimization
di: Schmidt, Carolin, et al.
Pubblicazione: (2024)
di: Schmidt, Carolin, et al.
Pubblicazione: (2024)
Operator Models for Continuous-Time Offline Reinforcement Learning
di: Hoischen, Nicolas, et al.
Pubblicazione: (2025)
di: Hoischen, Nicolas, et al.
Pubblicazione: (2025)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
di: Chen, Zixiang, et al.
Pubblicazione: (2025)
di: Chen, Zixiang, et al.
Pubblicazione: (2025)
Deflated Dynamics Value Iteration
di: Lee, Jongmin, et al.
Pubblicazione: (2024)
di: Lee, Jongmin, et al.
Pubblicazione: (2024)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
di: Di, Qiwei, et al.
Pubblicazione: (2024)
di: Di, Qiwei, et al.
Pubblicazione: (2024)
Rank-One Modified Value Iteration
di: Kolarijani, Arman Sharifi, et al.
Pubblicazione: (2025)
di: Kolarijani, Arman Sharifi, et al.
Pubblicazione: (2025)
Offline Reinforcement Learning via Linear-Programming with Error-Bound Induced Constraints
di: Ozdaglar, Asuman, et al.
Pubblicazione: (2022)
di: Ozdaglar, Asuman, et al.
Pubblicazione: (2022)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
di: Kou, Yiwen, et al.
Pubblicazione: (2024)
di: Kou, Yiwen, et al.
Pubblicazione: (2024)
Iteratively Reweighted Least Squares for Phase Unwrapping
di: Dubois-Taine, Benjamin, et al.
Pubblicazione: (2024)
di: Dubois-Taine, Benjamin, et al.
Pubblicazione: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
di: Zhou, Dongruo, et al.
Pubblicazione: (2018)
di: Zhou, Dongruo, et al.
Pubblicazione: (2018)
MARS: Unleashing the Power of Variance Reduction for Training Large Models
di: Yuan, Huizhuo, et al.
Pubblicazione: (2024)
di: Yuan, Huizhuo, et al.
Pubblicazione: (2024)
Computation of Least Trimmed Squares: A Branch-and-Bound framework with Hyperplane Arrangement Enhancements
di: Meng, Xiang, et al.
Pubblicazione: (2026)
di: Meng, Xiang, et al.
Pubblicazione: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
Towards Optimal Offline Reinforcement Learning
di: Li, Mengmeng, et al.
Pubblicazione: (2025)
di: Li, Mengmeng, et al.
Pubblicazione: (2025)
Sampling-based Safe Reinforcement Learning for Nonlinear Dynamical Systems
di: Suttle, Wesley A., et al.
Pubblicazione: (2024)
di: Suttle, Wesley A., et al.
Pubblicazione: (2024)
Value Mirror Descent for Reinforcement Learning
di: Jia, Zhichao, et al.
Pubblicazione: (2026)
di: Jia, Zhichao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
di: Zhao, Heyang, et al.
Pubblicazione: (2023) -
Reinforcement Learning from Human Feedback with Active Queries
di: Ji, Kaixuan, et al.
Pubblicazione: (2024) -
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
di: Zhang, Weitong, et al.
Pubblicazione: (2021) -
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
di: Di, Qiwei, et al.
Pubblicazione: (2023) -
Feel-Good Thompson Sampling for Contextual Dueling Bandits
di: Li, Xuheng, et al.
Pubblicazione: (2024)