4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Chaoyang, Mirzaei, Ashkan, Goel, Vidit, Menapace, Willi, Siarohin, Aliaksandr, Vinella, Avalon, Vasilkovsky, Michael, Skorokhodov, Ivan, Shakhrai, Vladislav, Korolev, Sergey, Tulyakov, Sergey, Wonka, Peter
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909657572835328
author Wang, Chaoyang
Mirzaei, Ashkan
Goel, Vidit
Menapace, Willi
Siarohin, Aliaksandr
Vinella, Avalon
Vasilkovsky, Michael
Skorokhodov, Ivan
Shakhrai, Vladislav
Korolev, Sergey
Tulyakov, Sergey
Wonka, Peter
author_facet Wang, Chaoyang
Mirzaei, Ashkan
Goel, Vidit
Menapace, Willi
Siarohin, Aliaksandr
Vinella, Avalon
Vasilkovsky, Michael
Skorokhodov, Ivan
Shakhrai, Vladislav
Korolev, Sergey
Tulyakov, Sergey
Wonka, Peter
contents We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze current 4D video diffusion architectures that perform spatial and temporal attention either sequentially or in parallel within a two-stream design. We highlight the limitations of existing approaches and introduce a novel fused architecture that performs spatial and temporal attention within a single layer. The key to our method is a sparse attention pattern, where tokens attend to others in the same frame, at the same timestamp, or from the same viewpoint. In the second part, we extend existing 3D reconstruction algorithms by introducing a Gaussian head, a camera token replacement algorithm, and additional dynamic layers and training. Overall, we establish a new state of the art for 4D generation, improving both visual quality and reconstruction capability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
Wang, Chaoyang
Mirzaei, Ashkan
Goel, Vidit
Menapace, Willi
Siarohin, Aliaksandr
Vinella, Avalon
Vasilkovsky, Michael
Skorokhodov, Ivan
Shakhrai, Vladislav
Korolev, Sergey
Tulyakov, Sergey
Wonka, Peter
Computer Vision and Pattern Recognition
We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze current 4D video diffusion architectures that perform spatial and temporal attention either sequentially or in parallel within a two-stream design. We highlight the limitations of existing approaches and introduce a novel fused architecture that performs spatial and temporal attention within a single layer. The key to our method is a sparse attention pattern, where tokens attend to others in the same frame, at the same timestamp, or from the same viewpoint. In the second part, we extend existing 3D reconstruction algorithms by introducing a Gaussian head, a camera token replacement algorithm, and additional dynamic layers and training. Overall, we establish a new state of the art for 4D generation, improving both visual quality and reconstruction capability.
title 4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18839