Physical Object Understanding with a Physically Controllable World Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Venkatesh, Rahul, Kotar, Klemen, Chen, Lilian Naing, Lee, Wanhee, Ancone, Gia, Kim, Seungwoo, Wheeler, Luca Thomas, Watrous, Jared, Chen, Honglin, Bear, Daniel, Stojanov, Stefan, Yamins, Daniel LK
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918532648796160
author Venkatesh, Rahul
Kotar, Klemen
Chen, Lilian Naing
Lee, Wanhee
Ancone, Gia
Kim, Seungwoo
Wheeler, Luca Thomas
Watrous, Jared
Chen, Honglin
Bear, Daniel
Stojanov, Stefan
Yamins, Daniel LK
author_facet Venkatesh, Rahul
Kotar, Klemen
Chen, Lilian Naing
Lee, Wanhee
Ancone, Gia
Kim, Seungwoo
Wheeler, Luca Thomas
Watrous, Jared
Chen, Honglin
Bear, Daniel
Stojanov, Stefan
Yamins, Daniel LK
contents A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solving these tasks requires world models capable of inferring distributional states of the world from partial observations - capabilities that current architectures do not provide. We introduce a new class of probabilistic world models that support estimation of the probability of any visual variable, such as appearance and dynamics, conditioned on any other variables. Here, we identify that these models can be trained efficiently with autoregressive sequence modeling, yielding world models from which rich object understanding emerges. First, we demonstrate that our model captures the physical laws governing how objects move by generating multiple plausible future states of the world through sequential inference. Then, by analyzing motion correlations across these futures, we extract objects and articulated object subparts. Having discovered these objects, we show that our world model can manipulate them in 3D. Finally, we demonstrate how physical relationships between objects can be computed from the world model, enabling applications such as Visual Jenga.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00439
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Physical Object Understanding with a Physically Controllable World Model
Venkatesh, Rahul
Kotar, Klemen
Chen, Lilian Naing
Lee, Wanhee
Ancone, Gia
Kim, Seungwoo
Wheeler, Luca Thomas
Watrous, Jared
Chen, Honglin
Bear, Daniel
Stojanov, Stefan
Yamins, Daniel LK
Computer Vision and Pattern Recognition
A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solving these tasks requires world models capable of inferring distributional states of the world from partial observations - capabilities that current architectures do not provide. We introduce a new class of probabilistic world models that support estimation of the probability of any visual variable, such as appearance and dynamics, conditioned on any other variables. Here, we identify that these models can be trained efficiently with autoregressive sequence modeling, yielding world models from which rich object understanding emerges. First, we demonstrate that our model captures the physical laws governing how objects move by generating multiple plausible future states of the world through sequential inference. Then, by analyzing motion correlations across these futures, we extract objects and articulated object subparts. Having discovered these objects, we show that our world model can manipulate them in 3D. Finally, we demonstrate how physical relationships between objects can be computed from the world model, enabling applications such as Visual Jenga.
title Physical Object Understanding with a Physically Controllable World Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.00439