Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Che, Fengdi, Xiao, Chenjun, Mei, Jincheng, Dai, Bo, Gummadi, Ramki, Ramirez, Oscar A, Harris, Christopher K, Mahmood, A. Rupam, Schuurmans, Dale |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Average-DICE: Stationary Distribution Correction by Regression
by: Che, Fengdi, et al.
Published: (2025)
by: Che, Fengdi, et al.
Published: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)
by: Mei, Jincheng, et al.
Published: (2024)
A Tutorial: An Intuitive Explanation of Offline Reinforcement Learning Theory
by: Che, Fengdi
Published: (2025)
by: Che, Fengdi
Published: (2025)
Satisficing Exploration for Deep Reinforcement Learning
by: Arumugam, Dilip, et al.
Published: (2024)
by: Arumugam, Dilip, et al.
Published: (2024)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
by: Dai, Bo, et al.
Published: (2026)
by: Dai, Bo, et al.
Published: (2026)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Autoregressive Large Language Models are Computationally Universal
by: Schuurmans, Dale, et al.
Published: (2024)
by: Schuurmans, Dale, et al.
Published: (2024)
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Learning Without Time-Based Embodiment Resets in Soft-Actor Critic
by: Farrahi, Homayoon, et al.
Published: (2025)
by: Farrahi, Homayoon, et al.
Published: (2025)
Spectral Representation-based Reinforcement Learning
by: Gao, Chenxiao, et al.
Published: (2025)
by: Gao, Chenxiao, et al.
Published: (2025)
Representation Learning via Non-Contrastive Mutual Information
by: Guo, Zhaohan Daniel, et al.
Published: (2025)
by: Guo, Zhaohan Daniel, et al.
Published: (2025)
Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
by: Lan, Qingfeng, et al.
Published: (2025)
by: Lan, Qingfeng, et al.
Published: (2025)
Streaming Deep Reinforcement Learning Finally Works
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
Primal-Dual Spectral Representation for Off-policy Evaluation
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Plastic Learning with Deep Fourier Features
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
Universal computation is intrinsic to language model decoding
by: Lewandowski, Alex, et al.
Published: (2026)
by: Lewandowski, Alex, et al.
Published: (2026)
Perovskite Solar Cell Stability Analysis Using Entropy‐Based Support Vector Machines Learning
by: Rupam Bhaduri, et al.
Published: (2024)
by: Rupam Bhaduri, et al.
Published: (2024)
Learning to Optimize for Reinforcement Learning
by: Lan, Qingfeng, et al.
Published: (2023)
by: Lan, Qingfeng, et al.
Published: (2023)
Weight Clipping for Deep Continual and Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
by: Ishfaq, Haque, et al.
Published: (2024)
by: Ishfaq, Haque, et al.
Published: (2024)
On the Benefits of Over-parameterization for Out-of-Distribution Generalization
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
Boosting Pruned Networks with Linear Over-parameterization
by: Qian, Yu, et al.
Published: (2022)
by: Qian, Yu, et al.
Published: (2022)
Toward Understanding In-context vs. In-weight Learning
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
by: Li, Chenjun
Published: (2026)
by: Li, Chenjun
Published: (2026)
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
Stochastic Trust-Region Methods for Over-parameterized Models
by: Yang, Aike, et al.
Published: (2026)
by: Yang, Aike, et al.
Published: (2026)
Directions of Curvature as an Explanation for Loss of Plasticity
by: Lewandowski, Alex, et al.
Published: (2023)
by: Lewandowski, Alex, et al.
Published: (2023)
Posterior Uncertainty for Targeted Parameters in Bayesian Bootstrap Procedures
by: Sabbagh, Magid, et al.
Published: (2026)
by: Sabbagh, Magid, et al.
Published: (2026)
A faster algorithm for Vertex Cover parameterized by solution size
by: Harris, David G., et al.
Published: (2022)
by: Harris, David G., et al.
Published: (2022)
A Validation Approach to Over-parameterized Matrix and Image Recovery
by: Ding, Lijun, et al.
Published: (2022)
by: Ding, Lijun, et al.
Published: (2022)
Estimation of Over-parameterized Models from an Auto-Modeling Perspective
by: Jiang, Yiran, et al.
Published: (2022)
by: Jiang, Yiran, et al.
Published: (2022)
Similar Items
-
Average-DICE: Stationary Distribution Correction by Regression
by: Che, Fengdi, et al.
Published: (2025) -
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023) -
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024) -
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025) -
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)