PuzzleJAX: A Benchmark for Reasoning and Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Earle, Sam, Todd, Graham, Li, Yuchen, Khalifa, Ahmed, Nasir, Muhammad Umair, Jiang, Zehua, Banburski-Fahey, Andrzej, Togelius, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScriptDoctor: Automatic Generation of PuzzleScript Games via Large Language Models and Tree Search
by: Earle, Sam, et al.
Published: (2025)
by: Earle, Sam, et al.
Published: (2025)
Missed Connections: Lateral Thinking Puzzles for Large Language Models
by: Todd, Graham, et al.
Published: (2024)
by: Todd, Graham, et al.
Published: (2024)
DreamGarden: A Designer Assistant for Growing Games from a Single Prompt
by: Earle, Sam, et al.
Published: (2024)
by: Earle, Sam, et al.
Published: (2024)
PCGRL+: Scaling, Control and Generalization in Reinforcement Learning Level Generators
by: Earle, Sam, et al.
Published: (2024)
by: Earle, Sam, et al.
Published: (2024)
Autoverse: An Evolvable Game Language for Learning Robust Embodied Agents
by: Earle, Sam, et al.
Published: (2024)
by: Earle, Sam, et al.
Published: (2024)
Video Game Level Design as a Multi-Agent Reinforcement Learning Problem
by: Earle, Sam, et al.
Published: (2025)
by: Earle, Sam, et al.
Published: (2025)
GameTraversalBenchmark: Evaluating Planning Abilities Of Large Language Models Through Traversing 2D Game Maps
by: Nasir, Muhammad Umair, et al.
Published: (2024)
by: Nasir, Muhammad Umair, et al.
Published: (2024)
Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting Games
by: Wang, Shouren, et al.
Published: (2025)
by: Wang, Shouren, et al.
Published: (2025)
Continuous Program Search
by: Siper, Matthew, et al.
Published: (2026)
by: Siper, Matthew, et al.
Published: (2026)
Making New Connections: LLMs as Puzzle Generators for The New York Times' Connections Word Game
by: Merino, Tim, et al.
Published: (2024)
by: Merino, Tim, et al.
Published: (2024)
DataDignity: Training Data Attribution for Large Language Models
by: Li, Xiaomin, et al.
Published: (2026)
by: Li, Xiaomin, et al.
Published: (2026)
LLMatic: Neural Architecture Search via Large Language Models and Quality Diversity Optimization
by: Nasir, Muhammad U., et al.
Published: (2023)
by: Nasir, Muhammad U., et al.
Published: (2023)
Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling
by: Xu, Weijia, et al.
Published: (2023)
by: Xu, Weijia, et al.
Published: (2023)
GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
by: Li, Yuchen, et al.
Published: (2025)
by: Li, Yuchen, et al.
Published: (2025)
Mortar: Evolving Mechanics for Automatic Game Design
by: Nasir, Muhammad U., et al.
Published: (2025)
by: Nasir, Muhammad U., et al.
Published: (2025)
PCGRLLM: Large Language Model-Driven Reward Design for Procedural Content Generation Reinforcement Learning
by: Baek, In-Chang, et al.
Published: (2025)
by: Baek, In-Chang, et al.
Published: (2025)
Learning Local Constraints for Reinforcement-Learned Content Generators
by: Bhaumik, Debosmita, et al.
Published: (2026)
by: Bhaumik, Debosmita, et al.
Published: (2026)
Large Language Models and Games: A Survey and Roadmap
by: Gallotta, Roberto, et al.
Published: (2024)
by: Gallotta, Roberto, et al.
Published: (2024)
Word2World: Generating Stories and Worlds through Large Language Models
by: Nasir, Muhammad U., et al.
Published: (2024)
by: Nasir, Muhammad U., et al.
Published: (2024)
Evolutionary Level Repair
by: Bhaumik, Debosmita, et al.
Published: (2025)
by: Bhaumik, Debosmita, et al.
Published: (2025)
Amorphous Fortress Online: Collaboratively Designing Open-Ended Multi-Agent AI and Game Environments
by: Charity, M, et al.
Published: (2025)
by: Charity, M, et al.
Published: (2025)
DreamCraft: Text-Guided Generation of Functional 3D Environments in Minecraft
by: Earle, Sam, et al.
Published: (2024)
by: Earle, Sam, et al.
Published: (2024)
The Procedural Content Generation Benchmark: An Open-source Testbed for Generative Challenges in Games
by: Khalifa, Ahmed, et al.
Published: (2025)
by: Khalifa, Ahmed, et al.
Published: (2025)
Goals as Reward-Producing Programs
by: Davidson, Guy, et al.
Published: (2024)
by: Davidson, Guy, et al.
Published: (2024)
Real-time Animation Generation and Control on Rigged Models via Large Language Models
by: Huang, Han, et al.
Published: (2023)
by: Huang, Han, et al.
Published: (2023)
A Markovian Framing of WaveFunctionCollapse for Procedurally Generating Aesthetically Complex Environments
by: Yiu, Franklin, et al.
Published: (2025)
by: Yiu, Franklin, et al.
Published: (2025)
Ludax: A GPU-Accelerated Domain Specific Language for Board Games
by: Todd, Graham, et al.
Published: (2025)
by: Todd, Graham, et al.
Published: (2025)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
by: Earle, Sam, et al.
Published: (2026)
by: Earle, Sam, et al.
Published: (2026)
All Stories Are One Story: Emotional Arc Guided Procedural Game Level Generation
by: Wen, Yunge, et al.
Published: (2025)
by: Wen, Yunge, et al.
Published: (2025)
Social Conjuring: Multi-User Runtime Collaboration with AI in Building Virtual 3D Worlds
by: Kobenova, Amina, et al.
Published: (2024)
by: Kobenova, Amina, et al.
Published: (2024)
GAVEL: Generating Games Via Evolution and Language Models
by: Todd, Graham, et al.
Published: (2024)
by: Todd, Graham, et al.
Published: (2024)
The Garden of Forking Paths: Narrative Arc-Conditioned Gameplay Planning
by: Wen, Yunge, et al.
Published: (2026)
by: Wen, Yunge, et al.
Published: (2026)
LLMR: Real-time Prompting of Interactive Worlds using Large Language Models
by: De La Torre, Fernanda, et al.
Published: (2023)
by: De La Torre, Fernanda, et al.
Published: (2023)
Learning to Reason via Program Generation, Emulation, and Search
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
by: Al-Khalifa, Shahad, et al.
Published: (2024)
by: Al-Khalifa, Shahad, et al.
Published: (2024)
Word2Minecraft: Generating 3D Game Levels through Large Language Models
by: Huang, Shuo, et al.
Published: (2025)
by: Huang, Shuo, et al.
Published: (2025)
A Dynamical Systems Framework for Reinforcement Learning Safety and Robustness Verification
by: Nasir, Ahmed, et al.
Published: (2025)
by: Nasir, Ahmed, et al.
Published: (2025)
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
by: Zhuang, Tianyi, et al.
Published: (2025)
by: Zhuang, Tianyi, et al.
Published: (2025)
minimax: Efficient Baselines for Autocurricula in JAX
by: Jiang, Minqi, et al.
Published: (2023)
by: Jiang, Minqi, et al.
Published: (2023)
Similar Items
-
ScriptDoctor: Automatic Generation of PuzzleScript Games via Large Language Models and Tree Search
by: Earle, Sam, et al.
Published: (2025) -
Missed Connections: Lateral Thinking Puzzles for Large Language Models
by: Todd, Graham, et al.
Published: (2024) -
DreamGarden: A Designer Assistant for Growing Games from a Single Prompt
by: Earle, Sam, et al.
Published: (2024) -
PCGRL+: Scaling, Control and Generalization in Reinforcement Learning Level Generators
by: Earle, Sam, et al.
Published: (2024) -
Autoverse: An Evolvable Game Language for Learning Robust Embodied Agents
by: Earle, Sam, et al.
Published: (2024)