Baba is LLM: Reasoning in a Game with Dynamic Rules
Fuente:
arXiv
Saved in:
| Main Authors: | van Wetten, Fien, Plaat, Aske, van Duijn, Max |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Research Re: search & Re-search
by: Plaat, Aske
Published: (2024)
by: Plaat, Aske
Published: (2024)
Agentic Large Language Models, a survey
by: Plaat, Aske, et al.
Published: (2025)
by: Plaat, Aske, et al.
Published: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
CoComposer: LLM Multi-agent Collaborative Music Composition
by: Xing, Peiwen, et al.
Published: (2025)
by: Xing, Peiwen, et al.
Published: (2025)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
by: Wong, Annie, et al.
Published: (2025)
by: Wong, Annie, et al.
Published: (2025)
Mirror Mode in Fire Emblem: Beating Players at their own Game with Imitation and Reinforcement Learning
by: Smid, Yanna Elizabeth, et al.
Published: (2025)
by: Smid, Yanna Elizabeth, et al.
Published: (2025)
State Design Matters: How Representations Shape Dynamic Reasoning in Large Language Models
by: Wong, Annie, et al.
Published: (2026)
by: Wong, Annie, et al.
Published: (2026)
Multi-Step Reasoning with Large Language Models, a Survey
by: Plaat, Aske, et al.
Published: (2024)
by: Plaat, Aske, et al.
Published: (2024)
Analysis of Bluffing by DQN and CFR in Leduc Hold'em Poker
by: Zaciragic, Tarik, et al.
Published: (2025)
by: Zaciragic, Tarik, et al.
Published: (2025)
Assessing Reproducibility in Evolutionary Computation: A Case Study using Human- and LLM-based Assessment
by: Da Ros, Francesca, et al.
Published: (2026)
by: Da Ros, Francesca, et al.
Published: (2026)
ACTIVA: Amortized Causal Effect Estimation via Transformer-based Variational Autoencoder
by: Sauter, Andreas, et al.
Published: (2025)
by: Sauter, Andreas, et al.
Published: (2025)
CausalPlayground: Addressing Data-Generation Requirements in Cutting-Edge Causality Research
by: Sauter, Andreas W M, et al.
Published: (2024)
by: Sauter, Andreas W M, et al.
Published: (2024)
Towards properly implementing Theory of Mind in AI systems: An account of four misconceptions
by: van der Meulen, Ramira, et al.
Published: (2025)
by: van der Meulen, Ramira, et al.
Published: (2025)
Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning
by: Spoor, Lindsay, et al.
Published: (2025)
by: Spoor, Lindsay, et al.
Published: (2025)
Reset-free Reinforcement Learning with World Models
by: Yang, Zhao, et al.
Published: (2024)
by: Yang, Zhao, et al.
Published: (2024)
Reasoning Promotes Robustness in Theory of Mind Tasks
by: de Haan, Ian B., et al.
Published: (2026)
by: de Haan, Ian B., et al.
Published: (2026)
Chargax: A JAX Accelerated EV Charging Simulator
by: Ponse, Koen, et al.
Published: (2025)
by: Ponse, Koen, et al.
Published: (2025)
Explicitly Disentangled Representations in Object-Centric Learning
by: Majellaro, Riccardo, et al.
Published: (2024)
by: Majellaro, Riccardo, et al.
Published: (2024)
Solving Deep Reinforcement Learning Tasks with Evolution Strategies and Linear Policy Networks
by: Wong, Annie, et al.
Published: (2024)
by: Wong, Annie, et al.
Published: (2024)
Guiding Skill Discovery with Foundation Models
by: Yang, Zhao, et al.
Published: (2025)
by: Yang, Zhao, et al.
Published: (2025)
A Benchmark Study of Deep Reinforcement Learning Algorithms for the Container Stowage Planning Problem
by: Huang, Yunqi, et al.
Published: (2025)
by: Huang, Yunqi, et al.
Published: (2025)
A Hybrid Intelligence Method for Argument Mining
by: van der Meer, Michiel, et al.
Published: (2024)
by: van der Meer, Michiel, et al.
Published: (2024)
CORE: Towards Scalable and Efficient Causal Discovery with Reinforcement Learning
by: Sauter, Andreas W. M., et al.
Published: (2024)
by: Sauter, Andreas W. M., et al.
Published: (2024)
Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models
by: Seo, SeungWon, et al.
Published: (2026)
by: Seo, SeungWon, et al.
Published: (2026)
Reinforcement Learning for Sustainable Energy: A Survey
by: Ponse, Koen, et al.
Published: (2024)
by: Ponse, Koen, et al.
Published: (2024)
Deliberative Dynamics and Value Alignment in LLM Debates
by: Sachdeva, Pratik S., et al.
Published: (2025)
by: Sachdeva, Pratik S., et al.
Published: (2025)
EduGym: An Environment and Notebook Suite for Reinforcement Learning Education
by: Moerland, Thomas M., et al.
Published: (2023)
by: Moerland, Thomas M., et al.
Published: (2023)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
by: Shi, Jiajun, et al.
Published: (2025)
by: Shi, Jiajun, et al.
Published: (2025)
Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?
by: Froma, Lennard C., et al.
Published: (2026)
by: Froma, Lennard C., et al.
Published: (2026)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
by: Gui, Jiayi, et al.
Published: (2024)
by: Gui, Jiayi, et al.
Published: (2024)
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
by: van der Maden, Willem, et al.
Published: (2026)
by: van der Maden, Willem, et al.
Published: (2026)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Discriminative Rule Learning for Outcome-Guided Process Model Discovery
by: Norouzifar, Ali, et al.
Published: (2025)
by: Norouzifar, Ali, et al.
Published: (2025)
GameArena: Evaluating LLM Reasoning through Live Computer Games
by: Hu, Lanxiang, et al.
Published: (2024)
by: Hu, Lanxiang, et al.
Published: (2024)
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
by: Xie, Tian, et al.
Published: (2025)
by: Xie, Tian, et al.
Published: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
by: Liu, Che, et al.
Published: (2025)
by: Liu, Che, et al.
Published: (2025)
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
by: Shu, Yubo, et al.
Published: (2025)
by: Shu, Yubo, et al.
Published: (2025)
Comparing Reinforcement Learning and Human Learning using the Game of Hidden Rules
by: Pulick, Eric, et al.
Published: (2023)
by: Pulick, Eric, et al.
Published: (2023)
Similar Items
-
Research Re: search & Re-search
by: Plaat, Aske
Published: (2024) -
Agentic Large Language Models, a survey
by: Plaat, Aske, et al.
Published: (2025) -
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
by: Chen, Xi, et al.
Published: (2025) -
CoComposer: LLM Multi-agent Collaborative Music Composition
by: Xing, Peiwen, et al.
Published: (2025) -
Reasoning Capabilities of Large Language Models on Dynamic Tasks
by: Wong, Annie, et al.
Published: (2025)