DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Röpke, Willem, Coward, Samuel, Lupu, Andrei, Foster, Thomas, Rocktäschel, Tim, Foerster, Jakob |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
JaxUED: A simple and useable UED library in Jax
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
Learning to Reason at the Frontier of Learnability
by: Foster, Thomas, et al.
Published: (2025)
by: Foster, Thomas, et al.
Published: (2025)
Open-Endedness is Essential for Artificial Superhuman Intelligence
by: Hughes, Edward, et al.
Published: (2024)
by: Hughes, Edward, et al.
Published: (2024)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
by: Matthews, Michael, et al.
Published: (2024)
by: Matthews, Michael, et al.
Published: (2024)
Behaviour Distillation
by: Lupu, Andrei, et al.
Published: (2024)
by: Lupu, Andrei, et al.
Published: (2024)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
by: Lupu, Andrei, et al.
Published: (2025)
by: Lupu, Andrei, et al.
Published: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
by: Cook, Jonathan, et al.
Published: (2025)
by: Cook, Jonathan, et al.
Published: (2025)
JaxLife: An Open-Ended Agentic Simulator
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Refining Minimax Regret for Unsupervised Environment Design
by: Beukman, Michael, et al.
Published: (2024)
by: Beukman, Michael, et al.
Published: (2024)
Imagined Autocurricula
by: Güzel, Ahmet H., et al.
Published: (2025)
by: Güzel, Ahmet H., et al.
Published: (2025)
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
by: Matthews, Michael, et al.
Published: (2024)
by: Matthews, Michael, et al.
Published: (2024)
Deep SPI: Safe Policy Improvement via World Models
by: Delgrange, Florent, et al.
Published: (2025)
by: Delgrange, Florent, et al.
Published: (2025)
Multi-Agent Diagnostics for Robustness via Illuminated Diversity
by: Samvelyan, Mikayel, et al.
Published: (2024)
by: Samvelyan, Mikayel, et al.
Published: (2024)
Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting Diversity
by: Costales, Robby, et al.
Published: (2024)
by: Costales, Robby, et al.
Published: (2024)
DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking
by: Hu, Tianyi, et al.
Published: (2026)
by: Hu, Tianyi, et al.
Published: (2026)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
by: Sims, Anya, et al.
Published: (2024)
by: Sims, Anya, et al.
Published: (2024)
Preference-Based Alignment of Discrete Diffusion Models
by: Borso, Umberto, et al.
Published: (2025)
by: Borso, Umberto, et al.
Published: (2025)
minimax: Efficient Baselines for Autocurricula in JAX
by: Jiang, Minqi, et al.
Published: (2023)
by: Jiang, Minqi, et al.
Published: (2023)
OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination
by: Gessler, Tobias, et al.
Published: (2025)
by: Gessler, Tobias, et al.
Published: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Evolving Many Worlds: Towards Open-Ended Discovery in Petri Dish NCA via Population-Based Training
by: Berdica, Uljad, et al.
Published: (2026)
by: Berdica, Uljad, et al.
Published: (2026)
The Autonomy-Alignment Problem in Open-Ended Learning Robots: Formalising the Purpose Framework
by: Baldassarre, Gianluca, et al.
Published: (2024)
by: Baldassarre, Gianluca, et al.
Published: (2024)
A Model-Based Solution to the Offline Multi-Agent Reinforcement Learning Coordination Problem
by: Barde, Paul, et al.
Published: (2023)
by: Barde, Paul, et al.
Published: (2023)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024)
by: Alberts, Lize, et al.
Published: (2024)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Scaling Opponent Shaping to High Dimensional Games
by: Khan, Akbir, et al.
Published: (2023)
by: Khan, Akbir, et al.
Published: (2023)
Goal-Conditioned Agents that Learn Everything All at Once
by: Matthews, Michael, et al.
Published: (2026)
by: Matthews, Michael, et al.
Published: (2026)
Intent Factored Generation: Unleashing the Diversity in Your Language Model
by: Ahmed, Eltayeb, et al.
Published: (2025)
by: Ahmed, Eltayeb, et al.
Published: (2025)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
by: Nasvytis, Linas, et al.
Published: (2024)
by: Nasvytis, Linas, et al.
Published: (2024)
Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2024)
by: Ma, Rachel, et al.
Published: (2024)
Select to Perfect: Imitating desired behavior from large multi-agent data
by: Franzmeyer, Tim, et al.
Published: (2024)
by: Franzmeyer, Tim, et al.
Published: (2024)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
by: Rutherford, Alexander, et al.
Published: (2023)
by: Rutherford, Alexander, et al.
Published: (2023)
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
by: Zhang, Mengyu, et al.
Published: (2025)
by: Zhang, Mengyu, et al.
Published: (2025)
Similar Items
-
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
by: Samvelyan, Mikayel, et al.
Published: (2024) -
JaxUED: A simple and useable UED library in Jax
by: Coward, Samuel, et al.
Published: (2024) -
Learning to Reason at the Frontier of Learnability
by: Foster, Thomas, et al.
Published: (2025) -
Open-Endedness is Essential for Artificial Superhuman Intelligence
by: Hughes, Edward, et al.
Published: (2024) -
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
by: Matthews, Michael, et al.
Published: (2024)