Saved in:
| Main Authors: | Güzel, Ahmet H., Jackson, Matthew Thomas, Liesen, Jarek Luca, Rocktäschel, Tim, Foerster, Jakob Nicolaus, Bogunovic, Ilija, Parker-Holder, Jack |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.13341 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
by: Güzel, Ahmet H., et al.
Published: (2025)
by: Güzel, Ahmet H., et al.
Published: (2025)
A Clean Slate for Offline Reinforcement Learning
by: Jackson, Matthew Thomas, et al.
Published: (2025)
by: Jackson, Matthew Thomas, et al.
Published: (2025)
minimax: Efficient Baselines for Autocurricula in JAX
by: Jiang, Minqi, et al.
Published: (2023)
by: Jiang, Minqi, et al.
Published: (2023)
Behaviour Distillation
by: Lupu, Andrei, et al.
Published: (2024)
by: Lupu, Andrei, et al.
Published: (2024)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
by: Güzel, Ahmet H., et al.
Published: (2026)
by: Güzel, Ahmet H., et al.
Published: (2026)
Multi-Agent Diagnostics for Robustness via Illuminated Diversity
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
by: Röpke, Willem, et al.
Published: (2026)
by: Röpke, Willem, et al.
Published: (2026)
Open-Endedness is Essential for Artificial Superhuman Intelligence
by: Hughes, Edward, et al.
Published: (2024)
by: Hughes, Edward, et al.
Published: (2024)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
by: Cook, Jonathan, et al.
Published: (2025)
by: Cook, Jonathan, et al.
Published: (2025)
Can Learned Optimization Make Reinforcement Learning Less Difficult?
by: Goldie, Alexander David, et al.
Published: (2024)
by: Goldie, Alexander David, et al.
Published: (2024)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
by: Paglieri, Davide, et al.
Published: (2025)
by: Paglieri, Davide, et al.
Published: (2025)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
Adversarially Robust Decision Transformer
by: Tang, Xiaohang, et al.
Published: (2024)
by: Tang, Xiaohang, et al.
Published: (2024)
Discovering Temporally-Aware Reinforcement Learning Algorithms
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
NAVIX: Scaling MiniGrid Environments with JAX
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
by: Ramesh, Shyam Sundhar, et al.
Published: (2023)
by: Ramesh, Shyam Sundhar, et al.
Published: (2023)
RSPO: Regularized Self-Play Alignment of Large Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
Discovering Minimal Reinforcement Learning Environments
by: Liesen, Jarek, et al.
Published: (2024)
by: Liesen, Jarek, et al.
Published: (2024)
Policy-Guided Diffusion
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
by: Xiong, Zheng, et al.
Published: (2025)
by: Xiong, Zheng, et al.
Published: (2025)
Preference-Based Alignment of Discrete Diffusion Models
by: Borso, Umberto, et al.
Published: (2025)
by: Borso, Umberto, et al.
Published: (2025)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
by: Ziomek, Juliusz, et al.
Published: (2026)
by: Ziomek, Juliusz, et al.
Published: (2026)
Robust Multi-Objective Controlled Decoding of Large Language Models
by: Son, Seongho, et al.
Published: (2025)
by: Son, Seongho, et al.
Published: (2025)
Goal-Conditioned Agents that Learn Everything All at Once
by: Matthews, Michael, et al.
Published: (2026)
by: Matthews, Michael, et al.
Published: (2026)
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
by: Tang, Xiaohang, et al.
Published: (2026)
by: Tang, Xiaohang, et al.
Published: (2026)
Investigating Non-Transitivity in LLM-as-a-Judge
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Evolution Strategies at the Hyperscale
by: Sarkar, Bidipta, et al.
Published: (2025)
by: Sarkar, Bidipta, et al.
Published: (2025)
JaxUED: A simple and useable UED library in Jax
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
How Should We Meta-Learn Reinforcement Learning Algorithms?
by: Goldie, Alexander David, et al.
Published: (2025)
by: Goldie, Alexander David, et al.
Published: (2025)
Meta-Learning Objectives for Preference Optimization
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
by: Nasvytis, Linas, et al.
Published: (2024)
by: Nasvytis, Linas, et al.
Published: (2024)
Select to Perfect: Imitating desired behavior from large multi-agent data
by: Franzmeyer, Tim, et al.
Published: (2024)
by: Franzmeyer, Tim, et al.
Published: (2024)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
by: Matthews, Michael, et al.
Published: (2024)
by: Matthews, Michael, et al.
Published: (2024)
Sample Efficient Preference Alignment in LLMs via Active Exploration
by: Mehta, Viraj, et al.
Published: (2023)
by: Mehta, Viraj, et al.
Published: (2023)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
by: Sims, Anya, et al.
Published: (2024)
by: Sims, Anya, et al.
Published: (2024)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
Similar Items
-
Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
by: Güzel, Ahmet H., et al.
Published: (2025) -
A Clean Slate for Offline Reinforcement Learning
by: Jackson, Matthew Thomas, et al.
Published: (2025) -
minimax: Efficient Baselines for Autocurricula in JAX
by: Jiang, Minqi, et al.
Published: (2023) -
Behaviour Distillation
by: Lupu, Andrei, et al.
Published: (2024) -
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
by: Paglieri, Davide, et al.
Published: (2024)