Discovering Minimal Reinforcement Learning Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Liesen, Jarek, Lu, Chris, Lupu, Andrei, Foerster, Jakob N., Sprekeler, Henning, Lange, Robert T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Behaviour Distillation
by: Lupu, Andrei, et al.
Published: (2024)
by: Lupu, Andrei, et al.
Published: (2024)
A Clean Slate for Offline Reinforcement Learning
by: Jackson, Matthew Thomas, et al.
Published: (2025)
by: Jackson, Matthew Thomas, et al.
Published: (2025)
Discovering Temporally-Aware Reinforcement Learning Algorithms
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
NAVIX: Scaling MiniGrid Environments with JAX
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
Discovering Preference Optimization Algorithms with and for Large Language Models
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Imagined Autocurricula
by: Güzel, Ahmet H., et al.
Published: (2025)
by: Güzel, Ahmet H., et al.
Published: (2025)
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
by: Röpke, Willem, et al.
Published: (2026)
by: Röpke, Willem, et al.
Published: (2026)
Recurrent Reinforcement Learning with Memoroids
by: Morad, Steven, et al.
Published: (2024)
by: Morad, Steven, et al.
Published: (2024)
Can Learned Optimization Make Reinforcement Learning Less Difficult?
by: Goldie, Alexander David, et al.
Published: (2024)
by: Goldie, Alexander David, et al.
Published: (2024)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
by: Sims, Anya, et al.
Published: (2024)
by: Sims, Anya, et al.
Published: (2024)
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
by: Ellis, Benjamin, et al.
Published: (2024)
by: Ellis, Benjamin, et al.
Published: (2024)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Boosting Fairness and Robustness in Over-the-Air Federated Learning
by: Oksuz, Halil Yigit, et al.
Published: (2024)
by: Oksuz, Halil Yigit, et al.
Published: (2024)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
by: Matthews, Michael, et al.
Published: (2024)
by: Matthews, Michael, et al.
Published: (2024)
Improving Regret Approximation for Unsupervised Dynamic Environment Generation
by: Mead, Harry, et al.
Published: (2026)
by: Mead, Harry, et al.
Published: (2026)
Abstraction for Offline Goal-Conditioned Reinforcement Learning
by: Wibault, Clarisse, et al.
Published: (2026)
by: Wibault, Clarisse, et al.
Published: (2026)
Bootstrapping Task Spaces for Self-Improvement
by: Jiang, Minqi, et al.
Published: (2025)
by: Jiang, Minqi, et al.
Published: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
by: Lupu, Andrei, et al.
Published: (2025)
by: Lupu, Andrei, et al.
Published: (2025)
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
by: Beukman, Michael, et al.
Published: (2026)
by: Beukman, Michael, et al.
Published: (2026)
EvIL: Evolution Strategies for Generalisable Imitation Learning
by: Sapora, Silvia, et al.
Published: (2024)
by: Sapora, Silvia, et al.
Published: (2024)
A Model-Based Solution to the Offline Multi-Agent Reinforcement Learning Coordination Problem
by: Barde, Paul, et al.
Published: (2023)
by: Barde, Paul, et al.
Published: (2023)
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
by: Yamada, Yutaro, et al.
Published: (2025)
by: Yamada, Yutaro, et al.
Published: (2025)
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
by: Ruhdorfer, Constantin, et al.
Published: (2025)
by: Ruhdorfer, Constantin, et al.
Published: (2025)
How Should We Meta-Learn Reinforcement Learning Algorithms?
by: Goldie, Alexander David, et al.
Published: (2025)
by: Goldie, Alexander David, et al.
Published: (2025)
Multi-Agent Craftax: Benchmarking Open-Ended Multi-Agent Reinforcement Learning at the Hyperscale
by: Omari, Bassel Al, et al.
Published: (2025)
by: Omari, Bassel Al, et al.
Published: (2025)
JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
by: Rutherford, Alexander, et al.
Published: (2023)
by: Rutherford, Alexander, et al.
Published: (2023)
SOReL and TOReL: Two Methods for Fully Offline Reinforcement Learning
by: Fellows, Mattie, et al.
Published: (2025)
by: Fellows, Mattie, et al.
Published: (2025)
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
by: Matthews, Michael, et al.
Published: (2024)
by: Matthews, Michael, et al.
Published: (2024)
An Optimisation Framework for Unsupervised Environment Design
by: Monette, Nathan, et al.
Published: (2025)
by: Monette, Nathan, et al.
Published: (2025)
Biology-inspired joint distribution neurons based on Hierarchical Correlation Reconstruction allowing for multidirectional propagation of values and densities
by: Duda, Jarek
Published: (2024)
by: Duda, Jarek
Published: (2024)
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
by: Nasvytis, Linas, et al.
Published: (2024)
by: Nasvytis, Linas, et al.
Published: (2024)
Using Constraints to Discover Sparse and Alternative Subgroup Descriptions
by: Bach, Jakob
Published: (2024)
by: Bach, Jakob
Published: (2024)
Adaptive stable distribution and Hurst exponent by method of moments moving estimator for nonstationary time series
by: Duda, Jarek
Published: (2025)
by: Duda, Jarek
Published: (2025)
Adaptive Student's t-distribution with method of moments moving estimator for nonstationary time series
by: Duda, Jarek
Published: (2023)
by: Duda, Jarek
Published: (2023)
Learning Multi-Agent Communication with Contrastive Learning
by: Lo, Yat Long, et al.
Published: (2023)
by: Lo, Yat Long, et al.
Published: (2023)
JaxUED: A simple and useable UED library in Jax
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
Refining Minimax Regret for Unsupervised Environment Design
by: Beukman, Michael, et al.
Published: (2024)
by: Beukman, Michael, et al.
Published: (2024)
ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive Advantages
by: Jesson, Andrew, et al.
Published: (2023)
by: Jesson, Andrew, et al.
Published: (2023)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
Higher order PCA-like rotation-invariant features for detailed shape descriptors modulo rotation
by: Duda, Jarek
Published: (2026)
by: Duda, Jarek
Published: (2026)
Similar Items
-
Behaviour Distillation
by: Lupu, Andrei, et al.
Published: (2024) -
A Clean Slate for Offline Reinforcement Learning
by: Jackson, Matthew Thomas, et al.
Published: (2025) -
Discovering Temporally-Aware Reinforcement Learning Algorithms
by: Jackson, Matthew Thomas, et al.
Published: (2024) -
NAVIX: Scaling MiniGrid Environments with JAX
by: Pignatelli, Eduardo, et al.
Published: (2024) -
Discovering Preference Optimization Algorithms with and for Large Language Models
by: Lu, Chris, et al.
Published: (2024)