SOReL and TOReL: Two Methods for Fully Offline Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Fellows, Mattie, Wibault, Clarisse, Berdica, Uljad, Forkel, Johannes, Osborne, Michael A., Foerster, Jakob N. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Abstraction for Offline Goal-Conditioned Reinforcement Learning
by: Wibault, Clarisse, et al.
Published: (2026)
by: Wibault, Clarisse, et al.
Published: (2026)
Learning to Reason at the Frontier of Learnability
by: Foster, Thomas, et al.
Published: (2025)
by: Foster, Thomas, et al.
Published: (2025)
A Clean Slate for Offline Reinforcement Learning
by: Jackson, Matthew Thomas, et al.
Published: (2025)
by: Jackson, Matthew Thomas, et al.
Published: (2025)
Evolving Many Worlds: Towards Open-Ended Discovery in Petri Dish NCA via Population-Based Training
by: Berdica, Uljad, et al.
Published: (2026)
by: Berdica, Uljad, et al.
Published: (2026)
Recurrent Structural Policy Gradient for Partially Observable Mean Field Games
by: Wibault, Clarisse, et al.
Published: (2026)
by: Wibault, Clarisse, et al.
Published: (2026)
Intent Factored Generation: Unleashing the Diversity in Your Language Model
by: Ahmed, Eltayeb, et al.
Published: (2025)
by: Ahmed, Eltayeb, et al.
Published: (2025)
Reinforcement Learning Controllers for Soft Robots using Learned Environments
by: Berdica, Uljad, et al.
Published: (2024)
by: Berdica, Uljad, et al.
Published: (2024)
Evolution Strategies at the Hyperscale
by: Sarkar, Bidipta, et al.
Published: (2025)
by: Sarkar, Bidipta, et al.
Published: (2025)
Refining Minimax Regret for Unsupervised Environment Design
by: Beukman, Michael, et al.
Published: (2024)
by: Beukman, Michael, et al.
Published: (2024)
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
by: Ruhdorfer, Constantin, et al.
Published: (2025)
by: Ruhdorfer, Constantin, et al.
Published: (2025)
Procedural Generation of Algorithm Discovery Tasks in Machine Learning
by: Goldie, Alexander D., et al.
Published: (2026)
by: Goldie, Alexander D., et al.
Published: (2026)
When Do We Need LLMs? A Diagnostic for Language-Driven Bandits
by: Berdica, Uljad, et al.
Published: (2026)
by: Berdica, Uljad, et al.
Published: (2026)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
by: Sims, Anya, et al.
Published: (2024)
by: Sims, Anya, et al.
Published: (2024)
A Model-Based Solution to the Offline Multi-Agent Reinforcement Learning Coordination Problem
by: Barde, Paul, et al.
Published: (2023)
by: Barde, Paul, et al.
Published: (2023)
Simplifying Deep Temporal Difference Learning
by: Gallici, Matteo, et al.
Published: (2024)
by: Gallici, Matteo, et al.
Published: (2024)
DITTO: Offline Imitation Learning with World Models
by: DeMoss, Branton, et al.
Published: (2023)
by: DeMoss, Branton, et al.
Published: (2023)
High entropy leads to symmetry equivariant policies in Dec-POMDPs
by: Forkel, Johannes, et al.
Published: (2025)
by: Forkel, Johannes, et al.
Published: (2025)
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
by: Nasvytis, Linas, et al.
Published: (2024)
by: Nasvytis, Linas, et al.
Published: (2024)
Expected Return Symmetries
by: Muglich, Darius, et al.
Published: (2025)
by: Muglich, Darius, et al.
Published: (2025)
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
by: Ellis, Benjamin, et al.
Published: (2024)
by: Ellis, Benjamin, et al.
Published: (2024)
Artificial Generational Intelligence: Cultural Accumulation in Reinforcement Learning
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
Recurrent Reinforcement Learning with Memoroids
by: Morad, Steven, et al.
Published: (2024)
by: Morad, Steven, et al.
Published: (2024)
AI & Human Co-Improvement for Safer Co-Superintelligence
by: Weston, Jason, et al.
Published: (2025)
by: Weston, Jason, et al.
Published: (2025)
How Should We Meta-Learn Reinforcement Learning Algorithms?
by: Goldie, Alexander David, et al.
Published: (2025)
by: Goldie, Alexander David, et al.
Published: (2025)
Learning Multi-Agent Communication with Contrastive Learning
by: Lo, Yat Long, et al.
Published: (2023)
by: Lo, Yat Long, et al.
Published: (2023)
Can Learned Optimization Make Reinforcement Learning Less Difficult?
by: Goldie, Alexander David, et al.
Published: (2024)
by: Goldie, Alexander David, et al.
Published: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
by: Ackermann, Johannes, et al.
Published: (2024)
by: Ackermann, Johannes, et al.
Published: (2024)
JaxUED: A simple and useable UED library in Jax
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
Ad-Hoc Human-AI Coordination Challenge
by: Dizdarević, Tin, et al.
Published: (2025)
by: Dizdarević, Tin, et al.
Published: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Offline Trajectory Optimization for Offline Reinforcement Learning
by: Zhao, Ziqi, et al.
Published: (2024)
by: Zhao, Ziqi, et al.
Published: (2024)
Discovering Temporally-Aware Reinforcement Learning Algorithms
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
Two-Step Offline Preference-Based Reinforcement Learning with Constrained Actions
by: Xu, Yinglun, et al.
Published: (2023)
by: Xu, Yinglun, et al.
Published: (2023)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
by: Matthews, Michael, et al.
Published: (2024)
by: Matthews, Michael, et al.
Published: (2024)
JaxLife: An Open-Ended Agentic Simulator
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Inflation Forecasting Post‐COVID‐19: Evidence From Germany
by: Tiphaine Wibault
Published: (2026)
by: Tiphaine Wibault
Published: (2026)
AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
Preference Elicitation for Offline Reinforcement Learning
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Offline Reinforcement Learning with Imbalanced Datasets
by: Jiang, Li, et al.
Published: (2023)
by: Jiang, Li, et al.
Published: (2023)
Similar Items
-
Abstraction for Offline Goal-Conditioned Reinforcement Learning
by: Wibault, Clarisse, et al.
Published: (2026) -
Learning to Reason at the Frontier of Learnability
by: Foster, Thomas, et al.
Published: (2025) -
A Clean Slate for Offline Reinforcement Learning
by: Jackson, Matthew Thomas, et al.
Published: (2025) -
Evolving Many Worlds: Towards Open-Ended Discovery in Petri Dish NCA via Population-Based Training
by: Berdica, Uljad, et al.
Published: (2026) -
Recurrent Structural Policy Gradient for Partially Observable Mean Field Games
by: Wibault, Clarisse, et al.
Published: (2026)