Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Zhaoyang, Xu, Canwen, Liu, Boyi, Wang, Yite, Han, Siwei, Yao, Zhewei, Yao, Huaxiu, He, Yuxiong
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916040195178496
author Wang, Zhaoyang
Xu, Canwen
Liu, Boyi
Wang, Yite
Han, Siwei
Yao, Zhewei
Yao, Huaxiu
He, Yuxiong
author_facet Wang, Zhaoyang
Xu, Canwen
Liu, Boyi
Wang, Yite
Han, Siwei
Yao, Zhewei
Yao, Huaxiu
He, Yuxiong
contents Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. However, scaling such agent training is limited by the lack of diverse and reliable environments. In this paper, we propose Agent World Model (AWM), a fully synthetic environment generation pipeline. Using this pipeline, we scale to 1,000 environments covering everyday scenarios, in which agents can interact with rich toolsets and obtain high-quality observations. Notably, these environments are code-driven and backed by databases, providing more reliable and consistent state transitions than environments simulated by LLMs. Moreover, they enable more efficient agent interaction compared with collecting trajectories from realistic environments. To demonstrate the effectiveness of this resource, we perform large-scale reinforcement learning for multi-turn tool-use agents. Thanks to the fully executable environments and accessible database states, we can also design reliable reward functions. Experiments on three benchmarks show that training exclusively in synthetic environments, rather than benchmark-specific ones, yields strong out-of-distribution generalization. The code is available at https://github.com/Snowflake-Labs/agent-world-model.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10090
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
Wang, Zhaoyang
Xu, Canwen
Liu, Boyi
Wang, Yite
Han, Siwei
Yao, Zhewei
Yao, Huaxiu
He, Yuxiong
Artificial Intelligence
Computation and Language
Machine Learning
Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. However, scaling such agent training is limited by the lack of diverse and reliable environments. In this paper, we propose Agent World Model (AWM), a fully synthetic environment generation pipeline. Using this pipeline, we scale to 1,000 environments covering everyday scenarios, in which agents can interact with rich toolsets and obtain high-quality observations. Notably, these environments are code-driven and backed by databases, providing more reliable and consistent state transitions than environments simulated by LLMs. Moreover, they enable more efficient agent interaction compared with collecting trajectories from realistic environments. To demonstrate the effectiveness of this resource, we perform large-scale reinforcement learning for multi-turn tool-use agents. Thanks to the fully executable environments and accessible database states, we can also design reliable reward functions. Experiments on three benchmarks show that training exclusively in synthetic environments, rather than benchmark-specific ones, yields strong out-of-distribution generalization. The code is available at https://github.com/Snowflake-Labs/agent-world-model.
title Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.10090