Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Fuente:
arXiv
Saved in:
| Main Authors: | Vasan, Gautham, Elsayed, Mohamed, Azimi, Alireza, He, Jiamin, Shariar, Fahim, Bellinger, Colin, White, Martha, Mahmood, A. Rupam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Versatile and Generalizable Manipulation via Goal-Conditioned Reinforcement Learning with Grounded Object Detection
by: Wang, Huiyi, et al.
Published: (2025)
by: Wang, Huiyi, et al.
Published: (2025)
Streaming Deep Reinforcement Learning Finally Works
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
General and Efficient Visual Goal-Conditioned Reinforcement Learning using Object-Agnostic Masks
by: Shahriar, Fahim, et al.
Published: (2025)
by: Shahriar, Fahim, et al.
Published: (2025)
Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
by: Lan, Qingfeng, et al.
Published: (2025)
by: Lan, Qingfeng, et al.
Published: (2025)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
Revisiting Sparse Rewards for Goal-Reaching Reinforcement Learning
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Weight Clipping for Deep Continual and Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Intentional Updates for Streaming Reinforcement Learning
by: Sharifnassab, Arsalan, et al.
Published: (2026)
by: Sharifnassab, Arsalan, et al.
Published: (2026)
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Learning Without Time-Based Embodiment Resets in Soft-Actor Critic
by: Farrahi, Homayoon, et al.
Published: (2025)
by: Farrahi, Homayoon, et al.
Published: (2025)
Extending Differential Temporal Difference Methods for Episodic Problems
by: De Asis, Kris, et al.
Published: (2026)
by: De Asis, Kris, et al.
Published: (2026)
Dynamic Observation Policies in Observation Cost-Sensitive Reinforcement Learning
by: Bellinger, Colin, et al.
Published: (2023)
by: Bellinger, Colin, et al.
Published: (2023)
Investigating the Interplay of Prioritized Replay and Generalization
by: Panahi, Parham Mohammad, et al.
Published: (2024)
by: Panahi, Parham Mohammad, et al.
Published: (2024)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
by: He, Jiamin, et al.
Published: (2026)
by: He, Jiamin, et al.
Published: (2026)
Adaptive Replay Buffer for Offline-to-Online Reinforcement Learning
by: Song, Chihyeon, et al.
Published: (2025)
by: Song, Chihyeon, et al.
Published: (2025)
Ring Artifact and Non-Uniformity Correction Method for Improving XACT Imaging
by: Eldib, Mohamed Elsayed
Published: (2024)
by: Eldib, Mohamed Elsayed
Published: (2024)
Enhancing Deep Deterministic Policy Gradients on Continuous Control Tasks with Decoupled Prioritized Experience Replay
by: Lorasdagi, Mehmet Efe, et al.
Published: (2025)
by: Lorasdagi, Mehmet Efe, et al.
Published: (2025)
Deep Reinforcement Learning with Gradient Eligibility Traces
by: Elelimy, Esraa, et al.
Published: (2025)
by: Elelimy, Esraa, et al.
Published: (2025)
Leveraging Complementary Embeddings for Replay Selection in Continual Learning with Small Buffers
by: Yanowsky, Danit, et al.
Published: (2026)
by: Yanowsky, Danit, et al.
Published: (2026)
Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL
by: Sandhu, Dillon, et al.
Published: (2026)
by: Sandhu, Dillon, et al.
Published: (2026)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
by: Corrado, Nicholas E., et al.
Published: (2023)
by: Corrado, Nicholas E., et al.
Published: (2023)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Grundbildung Medien im Studiengang Erwachsenenbildung
by: Bellinger, Franziska
Published: (2023)
by: Bellinger, Franziska
Published: (2023)
The Transformation from Microfilm to Digital Storage and Access.
by: Bellinger, Meg
Published: (1998)
by: Bellinger, Meg
Published: (1998)
Silence in the Quagmire: The Vietnam War in U.S. Comics. By Harriet, E. H. Earle Lincoln, NE: University of Nebraska Press, 2025. 198 Pages. $30.00 (paperback). ISBN: 978‐1‐4962‐4054‐5
by: Gwendolyn Bellinger
Published: (2026)
by: Gwendolyn Bellinger
Published: (2026)
Maintaining Plasticity in Deep Continual Learning
by: Dohare, Shibhansh, et al.
Published: (2023)
by: Dohare, Shibhansh, et al.
Published: (2023)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Prototype-Based Continual Learning with Label-free Replay Buffer and Cluster Preservation Loss
by: Aghasanli, Agil, et al.
Published: (2025)
by: Aghasanli, Agil, et al.
Published: (2025)
TEAL: New Selection Strategy for Small Buffers in Experience Replay Class Incremental Learning
by: Shaul-Ariel, Shahar, et al.
Published: (2024)
by: Shaul-Ariel, Shahar, et al.
Published: (2024)
Differentially Private Policy Gradient
by: Rio, Alexandre, et al.
Published: (2025)
by: Rio, Alexandre, et al.
Published: (2025)
On the Batch Size Selection in Stochastic Gradient Methods Using No-Replacement Sampling
by: Boresta, Marco, et al.
Published: (2025)
by: Boresta, Marco, et al.
Published: (2025)
VII.—On the correlation of the lower Lias at Barrow-on-Soar, in the Leicestershire, with the same strata in Warwickshire, Worcestershire, and Gloucestershire; and on the occurrence of the remains of insects at Barrow and in Yorkshire
by: Brodie, Peter Bellinger
Published: (1867)
by: Brodie, Peter Bellinger
Published: (1867)
Compute-Update Federated Learning: A Lattice Coding Approach Over-the-Air
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Formal Ethical Obligations in Reinforcement Learning Agents: Verification and Policy Updates
by: Shea-Blymyer, Colin, et al.
Published: (2024)
by: Shea-Blymyer, Colin, et al.
Published: (2024)
Algorithm-Relative Trajectory Valuation in Policy Gradient Control
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
Ground-state selection via nonlinear quantum dissipation
by: Ataei, Alireza, et al.
Published: (2026)
by: Ataei, Alireza, et al.
Published: (2026)
Parallel $k$d-tree with Batch Updates
by: Men, Ziyang, et al.
Published: (2024)
by: Men, Ziyang, et al.
Published: (2024)
Quantifying the Energy Floor: Direct Measurement and Replay Buffer Bias in SAC-Based HVAC Control on sbsim
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Similar Items
-
Versatile and Generalizable Manipulation via Goal-Conditioned Reinforcement Learning with Grounded Object Detection
by: Wang, Huiyi, et al.
Published: (2025) -
Streaming Deep Reinforcement Learning Finally Works
by: Elsayed, Mohamed, et al.
Published: (2024) -
General and Efficient Visual Goal-Conditioned Reinforcement Learning using Object-Agnostic Masks
by: Shahriar, Fahim, et al.
Published: (2025) -
Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
by: Lan, Qingfeng, et al.
Published: (2025) -
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)