Q-Learning with Fine-Grained Gap-Dependent Regret
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haochen, Zheng, Zhong, Xue, Lingzhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gap-Dependent Bounds for Federated $Q$-learning
by: Zhang, Haochen, et al.
Published: (2025)
by: Zhang, Haochen, et al.
Published: (2025)
Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition
by: Zheng, Zhong, et al.
Published: (2024)
by: Zheng, Zhong, et al.
Published: (2024)
Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning
by: Zhang, Haochen, et al.
Published: (2025)
by: Zhang, Haochen, et al.
Published: (2025)
Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost
by: Zheng, Zhong, et al.
Published: (2024)
by: Zheng, Zhong, et al.
Published: (2024)
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
Federated Q-Learning: Linear Regret Speedup with Low Communication Cost
by: Zheng, Zhong, et al.
Published: (2023)
by: Zheng, Zhong, et al.
Published: (2023)
Smoothed Robust Phase Retrieval
by: Zheng, Zhong, et al.
Published: (2024)
by: Zheng, Zhong, et al.
Published: (2024)
A New Inexact Proximal Linear Algorithm with Adaptive Stopping Criteria for Robust Phase Retrieval
by: Zheng, Zhong, et al.
Published: (2023)
by: Zheng, Zhong, et al.
Published: (2023)
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
by: Chen, Shulun, et al.
Published: (2025)
by: Chen, Shulun, et al.
Published: (2025)
A Unified Combination Framework for Dependent Tests with Applications to Microbiome Association Studies
by: Yu, Xiufan, et al.
Published: (2024)
by: Yu, Xiufan, et al.
Published: (2024)
Distributed Networked Multi-task Learning
by: Hong, Lingzhou, et al.
Published: (2024)
by: Hong, Lingzhou, et al.
Published: (2024)
PrunedLoRA: Robust Gradient-Based structured pruning for Low-rank Adaptation in Fine-tuning
by: Yu, Xin, et al.
Published: (2025)
by: Yu, Xin, et al.
Published: (2025)
A Copula Graphical Model for Multi-Attribute Data using Optimal Transport
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Revisiting Online Learning Approach to Inverse Linear Optimization: A Fenchel$-$Young Loss Perspective and Gap-Dependent Regret Analysis
by: Sakaue, Shinsaku, et al.
Published: (2025)
by: Sakaue, Shinsaku, et al.
Published: (2025)
AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating Projections
by: Yu, Xin, et al.
Published: (2025)
by: Yu, Xin, et al.
Published: (2025)
EXACT: Explicit Attribute-Guided Decoding-Time Personalization
by: Yu, Xin, et al.
Published: (2026)
by: Yu, Xin, et al.
Published: (2026)
Tight Gap-Dependent Memory-Regret Trade-Off for Single-Pass Streaming Stochastic Multi-Armed Bandits
by: Ye, Zichun, et al.
Published: (2025)
by: Ye, Zichun, et al.
Published: (2025)
No-Regret Linear Bandits under Gap-Adjusted Misspecification
by: Liu, Chong, et al.
Published: (2025)
by: Liu, Chong, et al.
Published: (2025)
Data-Dependent Regret Bounds for Constrained MABs
by: Genalti, Gianmarco, et al.
Published: (2025)
by: Genalti, Gianmarco, et al.
Published: (2025)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
Tackling Heavy-Tailed Rewards in Reinforcement Learning with Function Approximation: Minimax Optimal and Instance-Dependent Regret Bounds
by: Huang, Jiayi, et al.
Published: (2023)
by: Huang, Jiayi, et al.
Published: (2023)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization
by: Lei, Xing, et al.
Published: (2025)
by: Lei, Xing, et al.
Published: (2025)
Strongly Consistent Community Detection in Popularity Adjusted Block Models
by: Yuan, Quan, et al.
Published: (2025)
by: Yuan, Quan, et al.
Published: (2025)
Hypothesis Testing for High-Dimensional Matrix-Valued Data
by: Cui, Shijie, et al.
Published: (2024)
by: Cui, Shijie, et al.
Published: (2024)
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret
by: Zhong, Han, et al.
Published: (2023)
by: Zhong, Han, et al.
Published: (2023)
SGD with Dependent Data: Optimal Estimation, Regret, and Inference
by: Shen, Yinan, et al.
Published: (2026)
by: Shen, Yinan, et al.
Published: (2026)
Easy as ABCs: Unifying Boltzmann Q-Learning and Counterfactual Regret Minimization
by: D'Amico-Wong, Luca, et al.
Published: (2024)
by: D'Amico-Wong, Luca, et al.
Published: (2024)
Graph-Dependent Regret Bounds in Multi-Armed Bandits with Interference
by: Jamshidi, Fateme, et al.
Published: (2025)
by: Jamshidi, Fateme, et al.
Published: (2025)
Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
by: Di Gennaro, Federico, et al.
Published: (2025)
by: Di Gennaro, Federico, et al.
Published: (2025)
Structure-Preserving Nonlinear Sufficient Dimension Reduction for Tensors
by: Lin, Dianjun, et al.
Published: (2025)
by: Lin, Dianjun, et al.
Published: (2025)
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
by: Singh, Rahul, et al.
Published: (2026)
by: Singh, Rahul, et al.
Published: (2026)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Minimax Regret Learning for Data with Heterogeneous Subgroups
by: Mo, Weibin, et al.
Published: (2024)
by: Mo, Weibin, et al.
Published: (2024)
Doubly robust estimation of causal effects for random object outcomes with continuous treatments
by: Bhattacharjee, Satarupa, et al.
Published: (2025)
by: Bhattacharjee, Satarupa, et al.
Published: (2025)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
by: Yu, Xin, et al.
Published: (2026)
by: Yu, Xin, et al.
Published: (2026)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Understanding the Statistical Accuracy-Communication Trade-off in Personalized Federated Learning with Minimax Guarantees
by: Yu, Xin, et al.
Published: (2024)
by: Yu, Xin, et al.
Published: (2024)
Mind the Gap: Structure-Aware Consistency in Preference Learning
by: Mohri, Mehryar, et al.
Published: (2026)
by: Mohri, Mehryar, et al.
Published: (2026)
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Similar Items
-
Gap-Dependent Bounds for Federated $Q$-learning
by: Zhang, Haochen, et al.
Published: (2025) -
Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition
by: Zheng, Zhong, et al.
Published: (2024) -
Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning
by: Zhang, Haochen, et al.
Published: (2025) -
Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost
by: Zheng, Zhong, et al.
Published: (2024) -
Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation
by: Zhang, Haochen, et al.
Published: (2026)