Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Zhong, Zhang, Haochen, Xue, Lingzhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost
von: Zheng, Zhong, et al.
Veröffentlicht: (2024)
von: Zheng, Zhong, et al.
Veröffentlicht: (2024)
Gap-Dependent Bounds for Federated $Q$-learning
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
Q-Learning with Fine-Grained Gap-Dependent Regret
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
Federated Q-Learning: Linear Regret Speedup with Low Communication Cost
von: Zheng, Zhong, et al.
Veröffentlicht: (2023)
von: Zheng, Zhong, et al.
Veröffentlicht: (2023)
Smoothed Robust Phase Retrieval
von: Zheng, Zhong, et al.
Veröffentlicht: (2024)
von: Zheng, Zhong, et al.
Veröffentlicht: (2024)
A New Inexact Proximal Linear Algorithm with Adaptive Stopping Criteria for Robust Phase Retrieval
von: Zheng, Zhong, et al.
Veröffentlicht: (2023)
von: Zheng, Zhong, et al.
Veröffentlicht: (2023)
A Copula Graphical Model for Multi-Attribute Data using Optimal Transport
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
A Unified Combination Framework for Dependent Tests with Applications to Microbiome Association Studies
von: Yu, Xiufan, et al.
Veröffentlicht: (2024)
von: Yu, Xiufan, et al.
Veröffentlicht: (2024)
Distributed Networked Multi-task Learning
von: Hong, Lingzhou, et al.
Veröffentlicht: (2024)
von: Hong, Lingzhou, et al.
Veröffentlicht: (2024)
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating Projections
von: Yu, Xin, et al.
Veröffentlicht: (2025)
von: Yu, Xin, et al.
Veröffentlicht: (2025)
Mirror Descent Actor Critic via Bounded Advantage Learning
von: Iwaki, Ryo
Veröffentlicht: (2025)
von: Iwaki, Ryo
Veröffentlicht: (2025)
EXACT: Explicit Attribute-Guided Decoding-Time Personalization
von: Yu, Xin, et al.
Veröffentlicht: (2026)
von: Yu, Xin, et al.
Veröffentlicht: (2026)
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
von: Wiltzer, Harley, et al.
Veröffentlicht: (2024)
von: Wiltzer, Harley, et al.
Veröffentlicht: (2024)
Skill or Luck? Return Decomposition via Advantage Functions
von: Pan, Hsiao-Ru, et al.
Veröffentlicht: (2024)
von: Pan, Hsiao-Ru, et al.
Veröffentlicht: (2024)
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization
von: Lei, Xing, et al.
Veröffentlicht: (2025)
von: Lei, Xing, et al.
Veröffentlicht: (2025)
Hypothesis Testing for High-Dimensional Matrix-Valued Data
von: Cui, Shijie, et al.
Veröffentlicht: (2024)
von: Cui, Shijie, et al.
Veröffentlicht: (2024)
Strongly Consistent Community Detection in Popularity Adjusted Block Models
von: Yuan, Quan, et al.
Veröffentlicht: (2025)
von: Yuan, Quan, et al.
Veröffentlicht: (2025)
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
von: Nguyen, Tuan Ngo, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan Ngo, et al.
Veröffentlicht: (2024)
Support Basis: Fast Attention Beyond Bounded Entries
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2025)
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2025)
Boosting Soft Q-Learning by Bounding
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
Structure-Preserving Nonlinear Sufficient Dimension Reduction for Tensors
von: Lin, Dianjun, et al.
Veröffentlicht: (2025)
von: Lin, Dianjun, et al.
Veröffentlicht: (2025)
Doubly robust estimation of causal effects for random object outcomes with continuous treatments
von: Bhattacharjee, Satarupa, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Satarupa, et al.
Veröffentlicht: (2025)
PrunedLoRA: Robust Gradient-Based structured pruning for Low-rank Adaptation in Fine-tuning
von: Yu, Xin, et al.
Veröffentlicht: (2025)
von: Yu, Xin, et al.
Veröffentlicht: (2025)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
von: Yu, Xin, et al.
Veröffentlicht: (2026)
von: Yu, Xin, et al.
Veröffentlicht: (2026)
Generalization Bounds for Dependent Data using Online-to-Batch Conversion
von: Chatterjee, Sagnik, et al.
Veröffentlicht: (2024)
von: Chatterjee, Sagnik, et al.
Veröffentlicht: (2024)
Understanding the Statistical Accuracy-Communication Trade-off in Personalized Federated Learning with Minimax Guarantees
von: Yu, Xin, et al.
Veröffentlicht: (2024)
von: Yu, Xin, et al.
Veröffentlicht: (2024)
Mind the Gap: Structure-Aware Consistency in Preference Learning
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
Q-Learning with Shift-Aware Upper Confidence Bound in Non-Stationary Reinforcement Learning
von: Bui, Ha Manh, et al.
Veröffentlicht: (2025)
von: Bui, Ha Manh, et al.
Veröffentlicht: (2025)
Statistical Convergence Rates of Optimal Transport Map Estimation between General Distributions
von: Ding, Yizhe, et al.
Veröffentlicht: (2024)
von: Ding, Yizhe, et al.
Veröffentlicht: (2024)
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
von: Xue, Shuchen, et al.
Veröffentlicht: (2025)
von: Xue, Shuchen, et al.
Veröffentlicht: (2025)
Topology-Preserving Neural Operator Learning via Hodge Decomposition
von: Zheng, Dongzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Dongzhe, et al.
Veröffentlicht: (2026)
Tackling Heavy-Tailed Rewards in Reinforcement Learning with Function Approximation: Minimax Optimal and Instance-Dependent Regret Bounds
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
LAD: Learning Advantage Distribution for Reasoning
von: Li, Wendi, et al.
Veröffentlicht: (2026)
von: Li, Wendi, et al.
Veröffentlicht: (2026)
Path Learning with Trajectory Advantage Regression
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
Reference Neural Operators: Learning the Smooth Dependence of Solutions of PDEs on Geometric Deformations
von: Cheng, Ze, et al.
Veröffentlicht: (2024)
von: Cheng, Ze, et al.
Veröffentlicht: (2024)
Quantum-Informed Machine Learning for Predicting Spatiotemporal Chaos with Practical Quantum Advantage
von: Wang, Maida, et al.
Veröffentlicht: (2025)
von: Wang, Maida, et al.
Veröffentlicht: (2025)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost
von: Zheng, Zhong, et al.
Veröffentlicht: (2024) -
Gap-Dependent Bounds for Federated $Q$-learning
von: Zhang, Haochen, et al.
Veröffentlicht: (2025) -
Q-Learning with Fine-Grained Gap-Dependent Regret
von: Zhang, Haochen, et al.
Veröffentlicht: (2025) -
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
von: Zhang, Haochen, et al.
Veröffentlicht: (2026) -
Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)