The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mayor, Walter, Obando-Ceron, Johan, Courville, Aaron, Castro, Pablo Samuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2026)
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2026)
In value-based deep reinforcement learning, a pruned network is a good network
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2025)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2025)
Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
von: Tang, Hongyao, et al.
Veröffentlicht: (2025)
von: Tang, Hongyao, et al.
Veröffentlicht: (2025)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
von: Sokar, Ghada, et al.
Veröffentlicht: (2024)
von: Sokar, Ghada, et al.
Veröffentlicht: (2024)
The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
A Mechanistic Analysis of Looped Reasoning Language Models
von: Blayney, Hugh, et al.
Veröffentlicht: (2026)
von: Blayney, Hugh, et al.
Veröffentlicht: (2026)
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
Neuroplastic Expansion in Deep Reinforcement Learning
von: Liu, Jiashun, et al.
Veröffentlicht: (2024)
von: Liu, Jiashun, et al.
Veröffentlicht: (2024)
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2025)
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2025)
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
A Survey of State Representation Learning for Deep Reinforcement Learning
von: Echchahed, Ayoub, et al.
Veröffentlicht: (2025)
von: Echchahed, Ayoub, et al.
Veröffentlicht: (2025)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
von: Shah, Vedant, et al.
Veröffentlicht: (2025)
von: Shah, Vedant, et al.
Veröffentlicht: (2025)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
The Formalism-Implementation Gap in Reinforcement Learning Research
von: Castro, Pablo Samuel
Veröffentlicht: (2025)
von: Castro, Pablo Samuel
Veröffentlicht: (2025)
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
von: Lavoie, Samuel, et al.
Veröffentlicht: (2025)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Inventory Networks: Toward Reliable Policy Optimization
von: Alvo, Matias, et al.
Veröffentlicht: (2023)
von: Alvo, Matias, et al.
Veröffentlicht: (2023)
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
von: Huo, Yingxiao, et al.
Veröffentlicht: (2026)
von: Huo, Yingxiao, et al.
Veröffentlicht: (2026)
Solving Deep Reinforcement Learning Tasks with Evolution Strategies and Linear Policy Networks
von: Wong, Annie, et al.
Veröffentlicht: (2024)
von: Wong, Annie, et al.
Veröffentlicht: (2024)
Scaling Up Data Parallelism in Decentralized Deep Learning
von: Xie, Bing, et al.
Veröffentlicht: (2025)
von: Xie, Bing, et al.
Veröffentlicht: (2025)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
von: Liu, Xuefeng, et al.
Veröffentlicht: (2023)
von: Liu, Xuefeng, et al.
Veröffentlicht: (2023)
DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management
von: Xie, Yaqi, et al.
Veröffentlicht: (2026)
von: Xie, Yaqi, et al.
Veröffentlicht: (2026)
A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning
von: Juliani, Arthur, et al.
Veröffentlicht: (2024)
von: Juliani, Arthur, et al.
Veröffentlicht: (2024)
CALE: Continuous Arcade Learning Environment
von: Farebrother, Jesse, et al.
Veröffentlicht: (2024)
von: Farebrother, Jesse, et al.
Veröffentlicht: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
The Impact of Quantization and Pruning on Deep Reinforcement Learning Models
von: Lu, Heng, et al.
Veröffentlicht: (2024)
von: Lu, Heng, et al.
Veröffentlicht: (2024)
Multi-Task Reinforcement Learning Enables Parameter Scaling
von: McLean, Reginald, et al.
Veröffentlicht: (2025)
von: McLean, Reginald, et al.
Veröffentlicht: (2025)
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
Adaptive Data Exploitation in Deep Reinforcement Learning
von: Yuan, Mingqi, et al.
Veröffentlicht: (2025)
von: Yuan, Mingqi, et al.
Veröffentlicht: (2025)
ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2025)
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2025)
LOQA: Learning with Opponent Q-Learning Awareness
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
Efficient Deep Reinforcement Learning with Predictive Processing Proximal Policy Optimization
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2022)
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2022)
Optimal Policy Sparsification and Low Rank Decomposition for Deep Reinforcement Learning
von: Goddla, Vikram
Veröffentlicht: (2024)
von: Goddla, Vikram
Veröffentlicht: (2024)
Ähnliche Einträge
-
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2026) -
In value-based deep reinforcement learning, a pruned network is a good network
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024) -
Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2025) -
Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
von: Tang, Hongyao, et al.
Veröffentlicht: (2025) -
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)