Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ting-Hsuan, Chen, Ying-Huan, Tu, Tao, Lee, Jie-Ying, Wu, Cho-Ying, Lin, Fangzhou, Zhang, Hengyuan, Paz, David, Huang, Xinyu, Guo, Yuliang, Liu, Yu-Lun, Wang, Yue, Ren, Liu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913161236447232
author Chen, Ting-Hsuan
Chen, Ying-Huan
Tu, Tao
Lee, Jie-Ying
Wu, Cho-Ying
Lin, Fangzhou
Zhang, Hengyuan
Paz, David
Huang, Xinyu
Guo, Yuliang
Liu, Yu-Lun
Wang, Yue
Ren, Liu
author_facet Chen, Ting-Hsuan
Chen, Ying-Huan
Tu, Tao
Lee, Jie-Ying
Wu, Cho-Ying
Lin, Fangzhou
Zhang, Hengyuan
Paz, David
Huang, Xinyu
Guo, Yuliang
Liu, Yu-Lun
Wang, Yue
Ren, Liu
contents Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view trajectories, amplifying cross-view inconsistency and temporal drift. We argue that 360° video generation offers a natural solution: panoramic coverage simplifies trajectory design and provides a strong global context for maintaining coherence. We introduce Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion, a controllable 360° video generation framework that synthesizes high-fidelity videos from sparse 360° inputs. The key idea is an explicit 3D Cache, reconstructed from the input, which serves as a geometric scaffold for any user-defined camera path. This allows the diffusion model to focus on photorealistic texture refinement while the 3D Cache enforces global geometric consistency. Experiments show that Pantheon360 achieves superior visual quality and unmatched geometric coherence, enabling reliable and flexible 360° scene generation for downstream simulation and digital-twin applications.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25449
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion
Chen, Ting-Hsuan
Chen, Ying-Huan
Tu, Tao
Lee, Jie-Ying
Wu, Cho-Ying
Lin, Fangzhou
Zhang, Hengyuan
Paz, David
Huang, Xinyu
Guo, Yuliang
Liu, Yu-Lun
Wang, Yue
Ren, Liu
Computer Vision and Pattern Recognition
Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view trajectories, amplifying cross-view inconsistency and temporal drift. We argue that 360° video generation offers a natural solution: panoramic coverage simplifies trajectory design and provides a strong global context for maintaining coherence. We introduce Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion, a controllable 360° video generation framework that synthesizes high-fidelity videos from sparse 360° inputs. The key idea is an explicit 3D Cache, reconstructed from the input, which serves as a geometric scaffold for any user-defined camera path. This allows the diffusion model to focus on photorealistic texture refinement while the 3D Cache enforces global geometric consistency. Experiments show that Pantheon360 achieves superior visual quality and unmatched geometric coherence, enabling reliable and flexible 360° scene generation for downstream simulation and digital-twin applications.
title Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.25449