Computer-Using World Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guan, Yiming, Yu, Rui, Zhang, John, Wang, Lu, Zhang, Chaoyun, Li, Liqun, Qiao, Bo, Qin, Si, Huang, He, Yang, Fangkai, Zhao, Pu, Wutschitz, Lukas, Kessler, Samuel, Inan, Huseyin A, Sim, Robert, Rajmohan, Saravan, Lin, Qingwei, Zhang, Dongmei
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908840730034176
author Guan, Yiming
Yu, Rui
Zhang, John
Wang, Lu
Zhang, Chaoyun
Li, Liqun
Qiao, Bo
Qin, Si
Huang, He
Yang, Fangkai
Zhao, Pu
Wutschitz, Lukas
Kessler, Samuel
Inan, Huseyin A
Sim, Robert
Rajmohan, Saravan
Lin, Qingwei
Zhang, Dongmei
author_facet Guan, Yiming
Yu, Rui
Zhang, John
Wang, Lu
Zhang, Chaoyun
Li, Liqun
Qiao, Bo
Qin, Si
Huang, He
Yang, Fangkai
Zhao, Pu
Wutschitz, Lukas
Kessler, Samuel
Inan, Huseyin A
Sim, Robert
Rajmohan, Saravan
Lin, Qingwei
Zhang, Dongmei
contents Agents operating in complex software environments benefit from reasoning about the consequences of their actions, as even a single incorrect user interface (UI) operation can derail long, artifact-preserving workflows. This challenge is particularly acute for computer-using scenarios, where real execution does not support counterfactual exploration, making large-scale trial-and-error learning and planning impractical despite the environment being fully digital and deterministic. We introduce the Computer-Using World Model (CUWM), a world model for desktop software that predicts the next UI state given the current state and a candidate action. CUWM adopts a two-stage factorization of UI dynamics: it first predicts a textual description of agent-relevant state changes, and then realizes these changes visually to synthesize the next screenshot. CUWM is trained on offline UI transitions collected from agents interacting with real Microsoft Office applications, and further refined with a lightweight reinforcement learning stage that aligns textual transition predictions with the structural requirements of computer-using environments. We evaluate CUWM via test-time action search, where a frozen agent uses the world model to simulate and compare candidate actions before execution. Across a range of Office tasks, world-model-guided test-time scaling improves decision quality and execution robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17365
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Computer-Using World Model
Guan, Yiming
Yu, Rui
Zhang, John
Wang, Lu
Zhang, Chaoyun
Li, Liqun
Qiao, Bo
Qin, Si
Huang, He
Yang, Fangkai
Zhao, Pu
Wutschitz, Lukas
Kessler, Samuel
Inan, Huseyin A
Sim, Robert
Rajmohan, Saravan
Lin, Qingwei
Zhang, Dongmei
Software Engineering
Agents operating in complex software environments benefit from reasoning about the consequences of their actions, as even a single incorrect user interface (UI) operation can derail long, artifact-preserving workflows. This challenge is particularly acute for computer-using scenarios, where real execution does not support counterfactual exploration, making large-scale trial-and-error learning and planning impractical despite the environment being fully digital and deterministic. We introduce the Computer-Using World Model (CUWM), a world model for desktop software that predicts the next UI state given the current state and a candidate action. CUWM adopts a two-stage factorization of UI dynamics: it first predicts a textual description of agent-relevant state changes, and then realizes these changes visually to synthesize the next screenshot. CUWM is trained on offline UI transitions collected from agents interacting with real Microsoft Office applications, and further refined with a lightweight reinforcement learning stage that aligns textual transition predictions with the structural requirements of computer-using environments. We evaluate CUWM via test-time action search, where a frozen agent uses the world model to simulate and compare candidate actions before execution. Across a range of Office tasks, world-model-guided test-time scaling improves decision quality and execution robustness.
title Computer-Using World Model
topic Software Engineering
url https://arxiv.org/abs/2602.17365