SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917112489967616 |
|---|---|
| author | Li, Xinyi Xia, Zaishuo Lu, Weyl Hao, Chenjie Chen, Yubei |
| author_facet | Li, Xinyi Xia, Zaishuo Lu, Weyl Hao, Chenjie Chen, Yubei |
| contents | Current world models lack a unified and controlled setting for systematic evaluation, making it difficult to assess whether they truly capture the underlying rules that govern environment dynamics. In this work, we address this open challenge by introducing the SmallWorld Benchmark, a testbed designed to assess world model capability under isolated and precisely controlled dynamics without relying on handcrafted reward signals. Using this benchmark, we conduct comprehensive experiments in the fully observable state space on representative architectures including Recurrent State Space Model, Transformer, Diffusion model, and Neural ODE, examining their behavior across six distinct domains. The experimental results reveal how effectively these models capture environment structure and how their predictions deteriorate over extended rollouts, highlighting both the strengths and limitations of current modeling paradigms and offering insights into future improvement directions in representation learning and dynamics modeling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_23465 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments Li, Xinyi Xia, Zaishuo Lu, Weyl Hao, Chenjie Chen, Yubei Machine Learning Current world models lack a unified and controlled setting for systematic evaluation, making it difficult to assess whether they truly capture the underlying rules that govern environment dynamics. In this work, we address this open challenge by introducing the SmallWorld Benchmark, a testbed designed to assess world model capability under isolated and precisely controlled dynamics without relying on handcrafted reward signals. Using this benchmark, we conduct comprehensive experiments in the fully observable state space on representative architectures including Recurrent State Space Model, Transformer, Diffusion model, and Neural ODE, examining their behavior across six distinct domains. The experimental results reveal how effectively these models capture environment structure and how their predictions deteriorate over extended rollouts, highlighting both the strengths and limitations of current modeling paradigms and offering insights into future improvement directions in representation learning and dynamics modeling. |
| title | SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2511.23465 |