PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhan, Jiahao, Li, Zizhang, Yu, Hong-Xing, Wu, Jiajun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916029768138752
author Zhan, Jiahao
Li, Zizhang
Yu, Hong-Xing
Wu, Jiajun
author_facet Zhan, Jiahao
Li, Zizhang
Yu, Hong-Xing
Wu, Jiajun
contents We introduce PerpetualWonder, a hybrid generative simulator that enables long-horizon, action-conditioned 4D scene generation from a single image. Current works fail at this task because their physical state is decoupled from their visual representation, which prevents generative refinements to update the underlying physics for subsequent interactions. PerpetualWonder solves this by introducing the first true closed-loop system. It features a novel unified representation that creates a bidirectional link between the physical state and visual primitives, allowing generative refinements to correct both the dynamics and appearance. It also introduces a robust update mechanism that gathers supervision from multiple viewpoints to resolve optimization ambiguity. Experiments demonstrate that from a single image, PerpetualWonder can successfully simulate complex, multi-step interactions from long-horizon actions, maintaining physical plausibility and visual consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04876
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
Zhan, Jiahao
Li, Zizhang
Yu, Hong-Xing
Wu, Jiajun
Computer Vision and Pattern Recognition
We introduce PerpetualWonder, a hybrid generative simulator that enables long-horizon, action-conditioned 4D scene generation from a single image. Current works fail at this task because their physical state is decoupled from their visual representation, which prevents generative refinements to update the underlying physics for subsequent interactions. PerpetualWonder solves this by introducing the first true closed-loop system. It features a novel unified representation that creates a bidirectional link between the physical state and visual primitives, allowing generative refinements to correct both the dynamics and appearance. It also introduces a robust update mechanism that gathers supervision from multiple viewpoints to resolve optimization ambiguity. Experiments demonstrate that from a single image, PerpetualWonder can successfully simulate complex, multi-step interactions from long-horizon actions, maintaining physical plausibility and visual consistency.
title PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04876