Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peddi, Rohith, Saurabh, Shanmugam, Shravan, Pallapothula, Likhitha, Xiang, Yu, Singla, Parag, Gogate, Vibhav
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912965116035072
author Peddi, Rohith
Saurabh
Shanmugam, Shravan
Pallapothula, Likhitha
Xiang, Yu
Singla, Parag
Gogate, Vibhav
author_facet Peddi, Rohith
Saurabh
Shanmugam, Shravan
Pallapothula, Likhitha
Xiang, Yu
Singla, Parag
Gogate, Vibhav
contents Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason only about currently visible objects, discard entities upon occlusion, and operate in 2D. To address this, we first introduce ActionGenome4D, a dataset that upgrades Action Genome videos into 4D scenes via feed-forward 3D reconstruction, world-frame oriented bounding boxes for every object involved in actions, and dense relationship annotations including for objects that are temporarily unobserved due to occlusion or camera motion. Building on this data, we formalize World Scene Graph Generation (WSGG), the task of constructing a world scene graph at each timestamp that encompasses all interacting objects in the scene, both observed and unobserved. We then propose three complementary methods, each exploring a different inductive bias for reasoning about unobserved objects: PWG (Persistent World Graph), which implements object permanence via a zero-order feature buffer; MWAE (Masked World Auto-Encoder), which reframes unobserved-object reasoning as masked completion with cross-view associative retrieval; and 4DST (4D Scene Transformer), which replaces the static buffer with differentiable per-object temporal attention enriched by 3D motion and camera-pose features. We further design and evaluate the performance of strong open-source Vision-Language Models on the WSGG task via a suite of Graph RAG-based approaches, establishing baselines for unlocalized relationship prediction. WSGG thus advances video scene understanding toward world-centric, temporally persistent, and interpretable scene reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13185
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
Peddi, Rohith
Saurabh
Shanmugam, Shravan
Pallapothula, Likhitha
Xiang, Yu
Singla, Parag
Gogate, Vibhav
Computer Vision and Pattern Recognition
Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason only about currently visible objects, discard entities upon occlusion, and operate in 2D. To address this, we first introduce ActionGenome4D, a dataset that upgrades Action Genome videos into 4D scenes via feed-forward 3D reconstruction, world-frame oriented bounding boxes for every object involved in actions, and dense relationship annotations including for objects that are temporarily unobserved due to occlusion or camera motion. Building on this data, we formalize World Scene Graph Generation (WSGG), the task of constructing a world scene graph at each timestamp that encompasses all interacting objects in the scene, both observed and unobserved. We then propose three complementary methods, each exploring a different inductive bias for reasoning about unobserved objects: PWG (Persistent World Graph), which implements object permanence via a zero-order feature buffer; MWAE (Masked World Auto-Encoder), which reframes unobserved-object reasoning as masked completion with cross-view associative retrieval; and 4DST (4D Scene Transformer), which replaces the static buffer with differentiable per-object temporal attention enriched by 3D motion and camera-pose features. We further design and evaluate the performance of strong open-source Vision-Language Models on the WSGG task via a suite of Graph RAG-based approaches, establishing baselines for unlocalized relationship prediction. WSGG thus advances video scene understanding toward world-centric, temporally persistent, and interpretable scene reasoning.
title Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13185