Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Qiyue, Pi, Xinyu, Liu, Kevin, Chen, Junrong, Yang, Ruolan, Huang, Xinqi, Fang, Xinyu, Sun, Lu, Kishore, Gautham, Ai, Bo, Tao, Stone, Liu, Mengyang, Yang, Jiaxi, Lai, Chao-Jung, Jin, Chuanyang, Xiang, Jiannan, Huang, Benhao, Chen, Zeming, Danks, David, Su, Hao, Shu, Tianmin, Ma, Ziqiao, Qin, Lianhui, Hu, Zhiting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913914858504192
author Gao, Qiyue
Pi, Xinyu
Liu, Kevin
Chen, Junrong
Yang, Ruolan
Huang, Xinqi
Fang, Xinyu
Sun, Lu
Kishore, Gautham
Ai, Bo
Tao, Stone
Liu, Mengyang
Yang, Jiaxi
Lai, Chao-Jung
Jin, Chuanyang
Xiang, Jiannan
Huang, Benhao
Chen, Zeming
Danks, David
Su, Hao
Shu, Tianmin
Ma, Ziqiao
Qin, Lianhui
Hu, Zhiting
author_facet Gao, Qiyue
Pi, Xinyu
Liu, Kevin
Chen, Junrong
Yang, Ruolan
Huang, Xinqi
Fang, Xinyu
Sun, Lu
Kishore, Gautham
Ai, Bo
Tao, Stone
Liu, Mengyang
Yang, Jiaxi
Lai, Chao-Jung
Jin, Chuanyang
Xiang, Jiannan
Huang, Benhao
Chen, Zeming
Danks, David
Su, Hao
Shu, Tianmin
Ma, Ziqiao
Qin, Lianhui
Hu, Zhiting
contents Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Language Models (VLMs), such as OpenAI o3, GPT-4o and Gemini, exhibit potential as general-purpose WMs. While the latest studies have evaluated and shown limitations in specific capabilities such as visual understanding, a systematic evaluation of VLMs' fundamental WM abilities remains absent. Drawing on comparative psychology and cognitive science, we propose a two-stage framework that assesses Perception (visual, spatial, temporal, quantitative, and motion) and Prediction (mechanistic simulation, transitive inference, compositional inference) to provide an atomic evaluation of VLMs as WMs. Guided by this framework, we introduce WM-ABench, a large-scale benchmark comprising 23 fine-grained evaluation dimensions across 6 diverse simulated environments with controlled counterfactual simulations. Through 660 experiments on 15 latest commercial and open-source VLMs, we find that these models exhibit striking limitations in basic world modeling abilities. For instance, almost all models perform at near-random accuracy when distinguishing motion trajectories. Additionally, they lack disentangled understanding -- e.g., some models tend to believe blue objects move faster than green ones. More rich results and analyses reveal significant gaps between VLMs and human-level world modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
Gao, Qiyue
Pi, Xinyu
Liu, Kevin
Chen, Junrong
Yang, Ruolan
Huang, Xinqi
Fang, Xinyu
Sun, Lu
Kishore, Gautham
Ai, Bo
Tao, Stone
Liu, Mengyang
Yang, Jiaxi
Lai, Chao-Jung
Jin, Chuanyang
Xiang, Jiannan
Huang, Benhao
Chen, Zeming
Danks, David
Su, Hao
Shu, Tianmin
Ma, Ziqiao
Qin, Lianhui
Hu, Zhiting
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Language Models (VLMs), such as OpenAI o3, GPT-4o and Gemini, exhibit potential as general-purpose WMs. While the latest studies have evaluated and shown limitations in specific capabilities such as visual understanding, a systematic evaluation of VLMs' fundamental WM abilities remains absent. Drawing on comparative psychology and cognitive science, we propose a two-stage framework that assesses Perception (visual, spatial, temporal, quantitative, and motion) and Prediction (mechanistic simulation, transitive inference, compositional inference) to provide an atomic evaluation of VLMs as WMs. Guided by this framework, we introduce WM-ABench, a large-scale benchmark comprising 23 fine-grained evaluation dimensions across 6 diverse simulated environments with controlled counterfactual simulations. Through 660 experiments on 15 latest commercial and open-source VLMs, we find that these models exhibit striking limitations in basic world modeling abilities. For instance, almost all models perform at near-random accuracy when distinguishing motion trajectories. Additionally, they lack disentangled understanding -- e.g., some models tend to believe blue objects move faster than green ones. More rich results and analyses reveal significant gaps between VLMs and human-level world modeling.
title Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21876