Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | Beukman, Michael, Khetarpal, Khimya, Zheng, Zeyu, Dabney, Will, Foerster, Jakob, Dennis, Michael, Lyle, Clare |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Disentangling the Causes of Plasticity Loss in Neural Networks
di: Lyle, Clare, et al.
Pubblicazione: (2024)
di: Lyle, Clare, et al.
Pubblicazione: (2024)
Normalization and effective learning rates in reinforcement learning
di: Lyle, Clare, et al.
Pubblicazione: (2024)
di: Lyle, Clare, et al.
Pubblicazione: (2024)
JaxUED: A simple and useable UED library in Jax
di: Coward, Samuel, et al.
Pubblicazione: (2024)
di: Coward, Samuel, et al.
Pubblicazione: (2024)
Robust Intervention Learning from Emergency Stop Interventions
di: Pronovost, Ethan, et al.
Pubblicazione: (2026)
di: Pronovost, Ethan, et al.
Pubblicazione: (2026)
Refining Minimax Regret for Unsupervised Environment Design
di: Beukman, Michael, et al.
Pubblicazione: (2024)
di: Beukman, Michael, et al.
Pubblicazione: (2024)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
di: Matthews, Michael, et al.
Pubblicazione: (2024)
di: Matthews, Michael, et al.
Pubblicazione: (2024)
A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning
di: Khetarpal, Khimya, et al.
Pubblicazione: (2024)
di: Khetarpal, Khimya, et al.
Pubblicazione: (2024)
High entropy leads to symmetry equivariant policies in Dec-POMDPs
di: Forkel, Johannes, et al.
Pubblicazione: (2025)
di: Forkel, Johannes, et al.
Pubblicazione: (2025)
An Optimisation Framework for Unsupervised Environment Design
di: Monette, Nathan, et al.
Pubblicazione: (2025)
di: Monette, Nathan, et al.
Pubblicazione: (2025)
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
di: Matthews, Michael, et al.
Pubblicazione: (2024)
di: Matthews, Michael, et al.
Pubblicazione: (2024)
Goal-Conditioned Agents that Learn Everything All at Once
di: Matthews, Michael, et al.
Pubblicazione: (2026)
di: Matthews, Michael, et al.
Pubblicazione: (2026)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
di: Rutherford, Alexander, et al.
Pubblicazione: (2024)
di: Rutherford, Alexander, et al.
Pubblicazione: (2024)
Plasticity as the Mirror of Empowerment
di: Abel, David, et al.
Pubblicazione: (2025)
di: Abel, David, et al.
Pubblicazione: (2025)
Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning
di: Lawson, Daniel, et al.
Pubblicazione: (2025)
di: Lawson, Daniel, et al.
Pubblicazione: (2025)
Optimizing Return Distributions with Distributional Dynamic Programming
di: Pires, Bernardo Ávila, et al.
Pubblicazione: (2025)
di: Pires, Bernardo Ávila, et al.
Pubblicazione: (2025)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
di: Rowland, Mark, et al.
Pubblicazione: (2024)
di: Rowland, Mark, et al.
Pubblicazione: (2024)
Representation Learning via Non-Contrastive Mutual Information
di: Guo, Zhaohan Daniel, et al.
Pubblicazione: (2025)
di: Guo, Zhaohan Daniel, et al.
Pubblicazione: (2025)
Improving Regret Approximation for Unsupervised Dynamic Environment Generation
di: Mead, Harry, et al.
Pubblicazione: (2026)
di: Mead, Harry, et al.
Pubblicazione: (2026)
Discovering Minimal Reinforcement Learning Environments
di: Liesen, Jarek, et al.
Pubblicazione: (2024)
di: Liesen, Jarek, et al.
Pubblicazione: (2024)
Learning Multi-Agent Communication with Contrastive Learning
di: Lo, Yat Long, et al.
Pubblicazione: (2023)
di: Lo, Yat Long, et al.
Pubblicazione: (2023)
Affordances Enable Partial World Modeling with LLMs
di: Khetarpal, Khimya, et al.
Pubblicazione: (2026)
di: Khetarpal, Khimya, et al.
Pubblicazione: (2026)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
di: Obando-Ceron, Johan, et al.
Pubblicazione: (2024)
di: Obando-Ceron, Johan, et al.
Pubblicazione: (2024)
JaxLife: An Open-Ended Agentic Simulator
di: Lu, Chris, et al.
Pubblicazione: (2024)
di: Lu, Chris, et al.
Pubblicazione: (2024)
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
di: Ruhdorfer, Constantin, et al.
Pubblicazione: (2025)
di: Ruhdorfer, Constantin, et al.
Pubblicazione: (2025)
Multi-Agent Craftax: Benchmarking Open-Ended Multi-Agent Reinforcement Learning at the Hyperscale
di: Omari, Bassel Al, et al.
Pubblicazione: (2025)
di: Omari, Bassel Al, et al.
Pubblicazione: (2025)
Colored Noise in PPO: Improved Exploration and Performance through Correlated Action Sampling
di: Hollenstein, Jakob, et al.
Pubblicazione: (2023)
di: Hollenstein, Jakob, et al.
Pubblicazione: (2023)
Lifelong Reinforcement Learning via Neuromodulation
di: Lee, Sebastian, et al.
Pubblicazione: (2024)
di: Lee, Sebastian, et al.
Pubblicazione: (2024)
Edge Caching Optimization with PPO and Transfer Learning for Dynamic Environments
di: Niknia, Farnaz, et al.
Pubblicazione: (2024)
di: Niknia, Farnaz, et al.
Pubblicazione: (2024)
TabFlex: Scaling Tabular Learning to Millions with Linear Attention
di: Zeng, Yuchen, et al.
Pubblicazione: (2025)
di: Zeng, Yuchen, et al.
Pubblicazione: (2025)
RobocupGym: A challenging continuous control benchmark in Robocup
di: Beukman, Michael, et al.
Pubblicazione: (2024)
di: Beukman, Michael, et al.
Pubblicazione: (2024)
Step-Size Decay and Structural Stagnation in Greedy Sparse Learning
di: Berná, Pablo M.
Pubblicazione: (2026)
di: Berná, Pablo M.
Pubblicazione: (2026)
Directional-Clamp PPO
di: Karpel, Gilad, et al.
Pubblicazione: (2025)
di: Karpel, Gilad, et al.
Pubblicazione: (2025)
Mirror Learning: A Unifying Framework of Policy Optimisation
di: Kuba, Jakub Grudzien, et al.
Pubblicazione: (2022)
di: Kuba, Jakub Grudzien, et al.
Pubblicazione: (2022)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
di: Sims, Anya, et al.
Pubblicazione: (2024)
di: Sims, Anya, et al.
Pubblicazione: (2024)
SOReL and TOReL: Two Methods for Fully Offline Reinforcement Learning
di: Fellows, Mattie, et al.
Pubblicazione: (2025)
di: Fellows, Mattie, et al.
Pubblicazione: (2025)
Context Parallelism for Scalable Million-Token Inference
di: Yang, Amy, et al.
Pubblicazione: (2024)
di: Yang, Amy, et al.
Pubblicazione: (2024)
What Can Grokking Teach Us About Learning Under Nonstationarity?
di: Lyle, Clare, et al.
Pubblicazione: (2025)
di: Lyle, Clare, et al.
Pubblicazione: (2025)
Abstraction for Offline Goal-Conditioned Reinforcement Learning
di: Wibault, Clarisse, et al.
Pubblicazione: (2026)
di: Wibault, Clarisse, et al.
Pubblicazione: (2026)
Parallel Scaling Law for Language Models
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
di: Yan, Kaizhuo, et al.
Pubblicazione: (2025)
di: Yan, Kaizhuo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Disentangling the Causes of Plasticity Loss in Neural Networks
di: Lyle, Clare, et al.
Pubblicazione: (2024) -
Normalization and effective learning rates in reinforcement learning
di: Lyle, Clare, et al.
Pubblicazione: (2024) -
JaxUED: A simple and useable UED library in Jax
di: Coward, Samuel, et al.
Pubblicazione: (2024) -
Robust Intervention Learning from Emergency Stop Interventions
di: Pronovost, Ethan, et al.
Pubblicazione: (2026) -
Refining Minimax Regret for Unsupervised Environment Design
di: Beukman, Michael, et al.
Pubblicazione: (2024)