Towards Parameter-Free Temporal Difference Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yunxiang, Schmidt, Mark, Babanezhad, Reza, Vaswani, Sharan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025)
by: Vaswani, Sharan, et al.
Published: (2025)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
by: Asad, Reza, et al.
Published: (2025)
by: Asad, Reza, et al.
Published: (2025)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021)
by: Vaswani, Sharan, et al.
Published: (2021)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
by: Vaswani, Sharan, et al.
Published: (2026)
by: Vaswani, Sharan, et al.
Published: (2026)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)
by: Dang, Anh, et al.
Published: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
by: Fox, Curtis, et al.
Published: (2025)
by: Fox, Curtis, et al.
Published: (2025)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
From Inverse Optimization to Feasibility to ERM
by: Mishra, Saurabh, et al.
Published: (2024)
by: Mishra, Saurabh, et al.
Published: (2024)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
by: Liu, Xingtu, et al.
Published: (2025)
by: Liu, Xingtu, et al.
Published: (2025)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
by: Lu, Michael, et al.
Published: (2026)
by: Lu, Michael, et al.
Published: (2026)
Improving OOD Generalization of Pre-trained Encoders via Aligned Embedding-Space Ensembles
by: Peng, Shuman, et al.
Published: (2024)
by: Peng, Shuman, et al.
Published: (2024)
Preserving Plasticity in Continual Learning with Adaptive Linearity Injection
by: Rohani, Seyed Roozbeh Razavi, et al.
Published: (2025)
by: Rohani, Seyed Roozbeh Razavi, et al.
Published: (2025)
AltGDmin: Alternating GD and Minimization for Partly-Decoupled (Federated) Optimization
by: Vaswani, Namrata
Published: (2025)
by: Vaswani, Namrata
Published: (2025)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Fast and Sample Efficient Multi-Task Representation Learning in Stochastic Contextual Bandits
by: Lin, Jiabin, et al.
Published: (2024)
by: Lin, Jiabin, et al.
Published: (2024)
A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio
by: Shehper, Ali, et al.
Published: (2026)
by: Shehper, Ali, et al.
Published: (2026)
An Analysis of Quantile Temporal-Difference Learning
by: Rowland, Mark, et al.
Published: (2023)
by: Rowland, Mark, et al.
Published: (2023)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025)
by: Sahu, Sharan
Published: (2025)
Iterative Methods via Locally Evolving Set Process
by: Zhou, Baojian, et al.
Published: (2024)
by: Zhou, Baojian, et al.
Published: (2024)
Temporal Difference Learning with Constrained Initial Representations
by: Lyu, Jiafei, et al.
Published: (2026)
by: Lyu, Jiafei, et al.
Published: (2026)
CGCMA: Conditionally-Gated Cross-Modal Attention for Event-Conditioned Asynchronous Fusion
by: Guo, Yunxiang
Published: (2026)
by: Guo, Yunxiang
Published: (2026)
Simplifying Deep Temporal Difference Learning
by: Gallici, Matteo, et al.
Published: (2024)
by: Gallici, Matteo, et al.
Published: (2024)
On the Statistical Benefits of Temporal Difference Learning
by: Cheikhi, David, et al.
Published: (2023)
by: Cheikhi, David, et al.
Published: (2023)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Byzantine-Resilient Federated PCA and Low Rank Column-wise Sensing
by: Singh, Ankit Pratap, et al.
Published: (2023)
by: Singh, Ankit Pratap, et al.
Published: (2023)
Noisy Low Rank Column-wise Sensing
by: Singh, Ankit Pratap, et al.
Published: (2024)
by: Singh, Ankit Pratap, et al.
Published: (2024)
Efficient Federated Low Rank Matrix Completion
by: Abbasi, Ahmed Ali, et al.
Published: (2024)
by: Abbasi, Ahmed Ali, et al.
Published: (2024)
Enhancing Policy Gradient with the Polyak Step-Size Adaption
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
Statistical Inference for Temporal Difference Learning with Linear Function Approximation
by: Wu, Weichen, et al.
Published: (2024)
by: Wu, Weichen, et al.
Published: (2024)
Community-Aware Temporal Walks: Parameter-Free Representation Learning on Continuous-Time Dynamic Graphs
by: Yu, He, et al.
Published: (2025)
by: Yu, He, et al.
Published: (2025)
Discerning Temporal Difference Learning
by: Ma, Jianfei
Published: (2023)
by: Ma, Jianfei
Published: (2023)
Backstepping Temporal Difference Learning
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
Reinforcement Learning From State and Temporal Differences
by: Weaver, Lex, et al.
Published: (2025)
by: Weaver, Lex, et al.
Published: (2025)
New Versions of Gradient Temporal Difference Learning
by: Lee, Donghwan, et al.
Published: (2021)
by: Lee, Donghwan, et al.
Published: (2021)
Traffic-Aware Optimal Taxi Placement Using Graph Neural Network-Based Reinforcement Learning
by: Khetarpaul, Sonia, et al.
Published: (2026)
by: Khetarpaul, Sonia, et al.
Published: (2026)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Democratizing Signal Processing and Machine Learning: Math Learning Equity for Elementary and Middle School Students
by: Vaswani, Namrata, et al.
Published: (2024)
by: Vaswani, Namrata, et al.
Published: (2024)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
Similar Items
-
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025) -
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
by: Asad, Reza, et al.
Published: (2025) -
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
by: Vaswani, Sharan, et al.
Published: (2021) -
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
by: Vaswani, Sharan, et al.
Published: (2026) -
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)