Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yufeng, Cai, Qi, Yang, Zhuoran, Chen, Yongxin, Wang, Zhaoran |
|---|---|
| Format: | Preprint |
| Published: |
2020
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
by: Zhang, Yufeng, et al.
Published: (2021)
by: Zhang, Yufeng, et al.
Published: (2021)
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
by: Zhu, Yuchen, et al.
Published: (2024)
by: Zhu, Yuchen, et al.
Published: (2024)
Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
by: Yang, Zhuoran, et al.
Published: (2020)
by: Yang, Zhuoran, et al.
Published: (2020)
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
by: Cai, Qi, et al.
Published: (2022)
by: Cai, Qi, et al.
Published: (2022)
Provably Efficient Exploration in Policy Optimization
by: Cai, Qi, et al.
Published: (2019)
by: Cai, Qi, et al.
Published: (2019)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
A Generalized Sinkhorn Algorithm for Mean-Field Schrödinger Bridge
by: Eldesoukey, Asmaa, et al.
Published: (2026)
by: Eldesoukey, Asmaa, et al.
Published: (2026)
Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach
by: Qiu, Shuang, et al.
Published: (2022)
by: Qiu, Shuang, et al.
Published: (2022)
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
by: Wang, Lingxiao, et al.
Published: (2022)
by: Wang, Lingxiao, et al.
Published: (2022)
Analysis of Multiscale Reinforcement Q-Learning Algorithms for Mean Field Control Games
by: Angiuli, Andrea, et al.
Published: (2024)
by: Angiuli, Andrea, et al.
Published: (2024)
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
by: Shen, Han, et al.
Published: (2024)
by: Shen, Han, et al.
Published: (2024)
Learning to Stop: Deep Learning for Mean Field Optimal Stopping
by: Magnino, Lorenzo, et al.
Published: (2024)
by: Magnino, Lorenzo, et al.
Published: (2024)
Gauss-Newton Temporal Difference Learning with Nonlinear Function Approximation
by: Ke, Zhifa, et al.
Published: (2023)
by: Ke, Zhifa, et al.
Published: (2023)
Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
by: Fouque, Jean-Pierre, et al.
Published: (2025)
by: Fouque, Jean-Pierre, et al.
Published: (2025)
Structured Difference-of-Q via Orthogonal Learning
by: Cao, Defu, et al.
Published: (2024)
by: Cao, Defu, et al.
Published: (2024)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
by: Kim, Juno, et al.
Published: (2023)
by: Kim, Juno, et al.
Published: (2023)
Mean-Field Generalisation Bounds for Learning Controls in Stochastic Environments
by: Baros, Boris, et al.
Published: (2025)
by: Baros, Boris, et al.
Published: (2025)
Federated Temporal Difference Learning with Linear Function Approximation under Environmental Heterogeneity
by: Wang, Han, et al.
Published: (2023)
by: Wang, Han, et al.
Published: (2023)
Learning Surrogate Potential Mean Field Games via Gaussian Processes: A Data-Driven Approach to Ill-Posed Inverse Problems
by: Zhang, Jingguo, et al.
Published: (2025)
by: Zhang, Jingguo, et al.
Published: (2025)
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
by: Li, Wenye, et al.
Published: (2025)
by: Li, Wenye, et al.
Published: (2025)
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
by: Carmona, René, et al.
Published: (2019)
by: Carmona, René, et al.
Published: (2019)
Deep Reinforcement Learning for Infinite Horizon Mean Field Problems in Continuous Spaces
by: Angiuli, Andrea, et al.
Published: (2023)
by: Angiuli, Andrea, et al.
Published: (2023)
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
by: Nanda, Phalguni, et al.
Published: (2025)
by: Nanda, Phalguni, et al.
Published: (2025)
On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking
by: He, Jianliang, et al.
Published: (2026)
by: He, Jianliang, et al.
Published: (2026)
Inexact Column Generation for Bayesian Network Structure Learning via Difference-of-Submodular Optimization
by: Yang, Yiran, et al.
Published: (2025)
by: Yang, Yiran, et al.
Published: (2025)
Proximal Oracles for Optimization and Sampling
by: Liang, Jiaming, et al.
Published: (2024)
by: Liang, Jiaming, et al.
Published: (2024)
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
by: Wang, Shengbo
Published: (2026)
by: Wang, Shengbo
Published: (2026)
Data-Driven Adversarial Online Control for Unknown Linear Systems
by: Liu, Zishun, et al.
Published: (2023)
by: Liu, Zishun, et al.
Published: (2023)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
by: Mitra, Aritra, et al.
Published: (2023)
by: Mitra, Aritra, et al.
Published: (2023)
Novel clustered federated learning based on local loss
by: Gu, Endong, et al.
Published: (2024)
by: Gu, Endong, et al.
Published: (2024)
Mirror Mean-Field Langevin Dynamics
by: Gu, Anming, et al.
Published: (2025)
by: Gu, Anming, et al.
Published: (2025)
An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks
by: Ke, Zhifa, et al.
Published: (2024)
by: Ke, Zhifa, et al.
Published: (2024)
Universal Approximation Theorem for Deep Q-Learning via FBSDE System
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Major-Minor Mean Field Multi-Agent Reinforcement Learning
by: Cui, Kai, et al.
Published: (2023)
by: Cui, Kai, et al.
Published: (2023)
ADDQ: Adaptive Distributional Double Q-Learning
by: Döring, Leif, et al.
Published: (2025)
by: Döring, Leif, et al.
Published: (2025)
The Role of Target Update Frequencies in Q-Learning
by: Weissmann, Simon, et al.
Published: (2026)
by: Weissmann, Simon, et al.
Published: (2026)
Pointer Networks with Q-Learning for Combinatorial Optimization
by: Barro, Alessandro
Published: (2023)
by: Barro, Alessandro
Published: (2023)
Robust Q-Learning under Corrupted Rewards
by: Maity, Sreejeet, et al.
Published: (2024)
by: Maity, Sreejeet, et al.
Published: (2024)
Central Limit Theorems for Asynchronous Averaged Q-Learning
by: Liu, Xingtu
Published: (2025)
by: Liu, Xingtu
Published: (2025)
Similar Items
-
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
by: Zhang, Yufeng, et al.
Published: (2021) -
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
by: Zhu, Yuchen, et al.
Published: (2024) -
Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
by: Yang, Zhuoran, et al.
Published: (2020) -
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
by: Cai, Qi, et al.
Published: (2022) -
Provably Efficient Exploration in Policy Optimization
by: Cai, Qi, et al.
Published: (2019)