Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xinyu, Xie, Zixuan, Zhang, Shangtong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features
di: Xie, Zixuan, et al.
Pubblicazione: (2025)
di: Xie, Zixuan, et al.
Pubblicazione: (2025)
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
di: Wang, Jiuqi, et al.
Pubblicazione: (2024)
di: Wang, Jiuqi, et al.
Pubblicazione: (2024)
MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
Almost Sure Convergence Rates of Stochastic Approximation and Reinforcement Learning via a Poisson-Moreau Drift
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
Almost Sure Convergence Rates and Concentration of Stochastic Approximation and Reinforcement Learning with Markovian Noise
di: Qian, Xiaochi, et al.
Pubblicazione: (2024)
di: Qian, Xiaochi, et al.
Pubblicazione: (2024)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
di: Blaser, Ethan, et al.
Pubblicazione: (2026)
di: Blaser, Ethan, et al.
Pubblicazione: (2026)
Extensions of Robbins-Siegmund Theorem with Applications in Reinforcement Learning
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
di: Liu, Shuze Daniel, et al.
Pubblicazione: (2024)
di: Liu, Shuze Daniel, et al.
Pubblicazione: (2024)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
Counterfactual Explanations for Continuous Action Reinforcement Learning
di: Dong, Shuyang, et al.
Pubblicazione: (2025)
di: Dong, Shuyang, et al.
Pubblicazione: (2025)
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
di: Du, Ally Yalei, et al.
Pubblicazione: (2024)
di: Du, Ally Yalei, et al.
Pubblicazione: (2024)
Group Fairness in Multi-Task Reinforcement Learning
di: Song, Kefan, et al.
Pubblicazione: (2025)
di: Song, Kefan, et al.
Pubblicazione: (2025)
Convergent World Representations and Divergent Tasks
di: Park, Core Francisco
Pubblicazione: (2026)
di: Park, Core Francisco
Pubblicazione: (2026)
Boosting Soft Q-Learning by Bounding
di: Adamczyk, Jacob, et al.
Pubblicazione: (2024)
di: Adamczyk, Jacob, et al.
Pubblicazione: (2024)
Gradient Flow Drifting: Generative Modeling via Wasserstein Gradient Flows of KDE-Approximated Divergences
di: Cao, Jiarui, et al.
Pubblicazione: (2026)
di: Cao, Jiarui, et al.
Pubblicazione: (2026)
Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian Noise
di: Blaser, Ethan, et al.
Pubblicazione: (2024)
di: Blaser, Ethan, et al.
Pubblicazione: (2024)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
di: Fu, Deqing, et al.
Pubblicazione: (2023)
di: Fu, Deqing, et al.
Pubblicazione: (2023)
Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization
di: Tang, Cheng, et al.
Pubblicazione: (2024)
di: Tang, Cheng, et al.
Pubblicazione: (2024)
Experience Replay Addresses Loss of Plasticity in Continual Learning
di: Wang, Jiuqi, et al.
Pubblicazione: (2025)
di: Wang, Jiuqi, et al.
Pubblicazione: (2025)
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
di: Zhang, Kehao, et al.
Pubblicazione: (2026)
di: Zhang, Kehao, et al.
Pubblicazione: (2026)
Theoretical Analysis of Meta Reinforcement Learning: Generalization Bounds and Convergence Guarantees
di: Wang, Cangqing, et al.
Pubblicazione: (2024)
di: Wang, Cangqing, et al.
Pubblicazione: (2024)
Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning
di: Bozkurt, Alper Kamil, et al.
Pubblicazione: (2026)
di: Bozkurt, Alper Kamil, et al.
Pubblicazione: (2026)
On the Rate of Convergence of Kolmogorov-Arnold Network Regression Estimators
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
di: Dong, Zixuan, et al.
Pubblicazione: (2022)
di: Dong, Zixuan, et al.
Pubblicazione: (2022)
Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models
di: Cheng, Hailing, et al.
Pubblicazione: (2026)
di: Cheng, Hailing, et al.
Pubblicazione: (2026)
Learning Distinguishable Representations in Deep Q-Networks for Linear Transfer
di: Sathish, Sooraj, et al.
Pubblicazione: (2025)
di: Sathish, Sooraj, et al.
Pubblicazione: (2025)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
Mode-Aware Non-Linear Tucker Autoencoder for Tensor-based Unsupervised Learning
di: Zheng, Junjing, et al.
Pubblicazione: (2025)
di: Zheng, Junjing, et al.
Pubblicazione: (2025)
Toward a Unified Lyapunov-Certified ODE Convergence Analysis of Smooth Q-Learning with p-Norms
di: Lee, Donghwan, et al.
Pubblicazione: (2024)
di: Lee, Donghwan, et al.
Pubblicazione: (2024)
Convergent Linear Representations of Emergent Misalignment
di: Soligo, Anna, et al.
Pubblicazione: (2025)
di: Soligo, Anna, et al.
Pubblicazione: (2025)
Linear Attention is Enough in Spatial-Temporal Forecasting
di: Ning, Xinyu
Pubblicazione: (2024)
di: Ning, Xinyu
Pubblicazione: (2024)
Graph Edit Distance with General Costs Using Neural Set Divergence
di: Jain, Eeshaan, et al.
Pubblicazione: (2024)
di: Jain, Eeshaan, et al.
Pubblicazione: (2024)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
di: Defazio, Aaron, et al.
Pubblicazione: (2023)
di: Defazio, Aaron, et al.
Pubblicazione: (2023)
Expert Divergence Learning for MoE-based Language Models
di: Li, Jiaang, et al.
Pubblicazione: (2026)
di: Li, Jiaang, et al.
Pubblicazione: (2026)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
di: Nguyen, Tien-Phat, et al.
Pubblicazione: (2026)
di: Nguyen, Tien-Phat, et al.
Pubblicazione: (2026)
Tight Lower Bounds and Improved Convergence in Performative Prediction
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
Efficient Representations for High-Cardinality Categorical Variables in Machine Learning
di: Liang, Zixuan
Pubblicazione: (2025)
di: Liang, Zixuan
Pubblicazione: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features
di: Xie, Zixuan, et al.
Pubblicazione: (2025) -
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
di: Wang, Jiuqi, et al.
Pubblicazione: (2024) -
MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries
di: Xie, Zixuan, et al.
Pubblicazione: (2026) -
Almost Sure Convergence Rates of Stochastic Approximation and Reinforcement Learning via a Poisson-Moreau Drift
di: Liu, Xinyu, et al.
Pubblicazione: (2026) -
Almost Sure Convergence Rates and Concentration of Stochastic Approximation and Reinforcement Learning with Markovian Noise
di: Qian, Xiaochi, et al.
Pubblicazione: (2024)