WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Upadhyay, Rishi, Zhang, Howard, Solomon, Jim, Agrawal, Ayush, Boreddy, Pranay, Narayana, Shruti Satya, Ba, Yunhao, Wong, Alex, de Melo, Celso M, Kadambi, Achuta
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917231039873024
author Upadhyay, Rishi
Zhang, Howard
Solomon, Jim
Agrawal, Ayush
Boreddy, Pranay
Narayana, Shruti Satya
Ba, Yunhao
Wong, Alex
de Melo, Celso M
Kadambi, Achuta
author_facet Upadhyay, Rishi
Zhang, Howard
Solomon, Jim
Agrawal, Ayush
Boreddy, Pranay
Narayana, Shruti Satya
Ba, Yunhao
Wong, Alex
de Melo, Celso M
Kadambi, Achuta
contents Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must exhibit high physical fidelity, accurately simulating real-world dynamics. Existing physics-based video benchmarks, however, suffer from entanglement, where a single test simultaneously evaluates multiple physical laws and concepts, fundamentally limiting their diagnostic capability. We introduce WorldBench, a novel video-based benchmark specifically designed for concept-specific, disentangled evaluation, allowing us to rigorously isolate and assess understanding of a single physical concept or law at a time. To make WorldBench comprehensive, we design benchmarks at two different levels: 1) an evaluation of intuitive physical understanding with concepts such as object permanence or scale/perspective, and 2) an evaluation of low-level physical constants and material properties such as friction coefficients or fluid viscosity. When SOTA video-based world models are evaluated on WorldBench, we find specific patterns of failure in particular physics concepts, with all tested models lacking the physical consistency required to generate reliable real-world interactions. Through its concept-specific evaluation, WorldBench offers a more nuanced and scalable framework for rigorously evaluating the physical reasoning capabilities of video generation and world models, paving the way for more robust and generalizable world-model-driven learning.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21282
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
Upadhyay, Rishi
Zhang, Howard
Solomon, Jim
Agrawal, Ayush
Boreddy, Pranay
Narayana, Shruti Satya
Ba, Yunhao
Wong, Alex
de Melo, Celso M
Kadambi, Achuta
Computer Vision and Pattern Recognition
Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must exhibit high physical fidelity, accurately simulating real-world dynamics. Existing physics-based video benchmarks, however, suffer from entanglement, where a single test simultaneously evaluates multiple physical laws and concepts, fundamentally limiting their diagnostic capability. We introduce WorldBench, a novel video-based benchmark specifically designed for concept-specific, disentangled evaluation, allowing us to rigorously isolate and assess understanding of a single physical concept or law at a time. To make WorldBench comprehensive, we design benchmarks at two different levels: 1) an evaluation of intuitive physical understanding with concepts such as object permanence or scale/perspective, and 2) an evaluation of low-level physical constants and material properties such as friction coefficients or fluid viscosity. When SOTA video-based world models are evaluated on WorldBench, we find specific patterns of failure in particular physics concepts, with all tested models lacking the physical consistency required to generate reliable real-world interactions. Through its concept-specific evaluation, WorldBench offers a more nuanced and scalable framework for rigorously evaluating the physical reasoning capabilities of video generation and world models, paving the way for more robust and generalizable world-model-driven learning.
title WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.21282