Current Agents Fail to Leverage World Model as Tool for Foresight

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Cheng, Acikgoz, Emre Can, Li, Bingxuan, Chen, Xiusi, Zhang, Yuji, He, Bingxiang, Luo, Qinyu, Hakkani-Tür, Dilek, Tur, Gokhan, Li, Yunzhu, Ji, Heng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909984212647936
author Qian, Cheng
Acikgoz, Emre Can
Li, Bingxuan
Chen, Xiusi
Zhang, Yuji
He, Bingxiang
Luo, Qinyu
Hakkani-Tür, Dilek
Tur, Gokhan
Li, Yunzhu
Ji, Heng
author_facet Qian, Cheng
Acikgoz, Emre Can
Li, Bingxuan
Chen, Xiusi
Zhang, Yuji
He, Bingxiang
Luo, Qinyu
Hakkani-Tür, Dilek
Tur, Gokhan
Li, Yunzhu
Ji, Heng
contents Agents built on vision-language models increasingly face tasks that demand anticipating future states rather than relying on short-horizon reasoning. Generative world models offer a promising remedy: agents could use them as external simulators to foresee outcomes before acting. This paper empirically examines whether current agents can leverage such world models as tools to enhance their cognition. Across diverse agentic and visual question answering tasks, we observe that some agents rarely invoke simulation (fewer than 1%), frequently misuse predicted rollouts (approximately 15%), and often exhibit inconsistent or even degraded performance (up to 5%) when simulation is available or enforced. Attribution analysis further indicates that the primary bottleneck lies in the agents' capacity to decide when to simulate, how to interpret predicted outcomes, and how to integrate foresight into downstream reasoning. These findings underscore the need for mechanisms that foster calibrated, strategic interaction with world models, paving the way toward more reliable anticipatory cognition in future agent systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03905
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Current Agents Fail to Leverage World Model as Tool for Foresight
Qian, Cheng
Acikgoz, Emre Can
Li, Bingxuan
Chen, Xiusi
Zhang, Yuji
He, Bingxiang
Luo, Qinyu
Hakkani-Tür, Dilek
Tur, Gokhan
Li, Yunzhu
Ji, Heng
Artificial Intelligence
Computation and Language
Machine Learning
Agents built on vision-language models increasingly face tasks that demand anticipating future states rather than relying on short-horizon reasoning. Generative world models offer a promising remedy: agents could use them as external simulators to foresee outcomes before acting. This paper empirically examines whether current agents can leverage such world models as tools to enhance their cognition. Across diverse agentic and visual question answering tasks, we observe that some agents rarely invoke simulation (fewer than 1%), frequently misuse predicted rollouts (approximately 15%), and often exhibit inconsistent or even degraded performance (up to 5%) when simulation is available or enforced. Attribution analysis further indicates that the primary bottleneck lies in the agents' capacity to decide when to simulate, how to interpret predicted outcomes, and how to integrate foresight into downstream reasoning. These findings underscore the need for mechanisms that foster calibrated, strategic interaction with world models, paving the way toward more reliable anticipatory cognition in future agent systems.
title Current Agents Fail to Leverage World Model as Tool for Foresight
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.03905