View-Consistent Diffusion Representations for 3D-Consistent Video Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Danier, Duolikun, Gao, Ge, McDonagh, Steven, Li, Changjian, Bilen, Hakan, Mac Aodha, Oisin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909920638533632
author Danier, Duolikun
Gao, Ge
McDonagh, Steven
Li, Changjian
Bilen, Hakan
Mac Aodha, Oisin
author_facet Danier, Duolikun
Gao, Ge
McDonagh, Steven
Li, Changjian
Bilen, Hakan
Mac Aodha, Oisin
contents Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D inconsistencies, e.g., objects and structures deforming under changes in camera pose, which can undermine user experience and simulation fidelity. Motivated by recent findings on representation alignment for diffusion models, we hypothesize that improving the multi-view consistency of video diffusion representations will yield more 3D-consistent video generation. Through detailed analysis on multiple recent camera-controlled video diffusion models we reveal strong correlations between 3D-consistent representations and videos. We also propose ViCoDR, a new approach for improving the 3D consistency of video models by learning multi-view consistent diffusion representations. We evaluate ViCoDR on camera controlled image-to-video, text-to-video, and multi-view generation models, demonstrating significant improvements in the 3D consistency of the generated videos. Project page: https://danier97.github.io/ViCoDR.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle View-Consistent Diffusion Representations for 3D-Consistent Video Generation
Danier, Duolikun
Gao, Ge
McDonagh, Steven
Li, Changjian
Bilen, Hakan
Mac Aodha, Oisin
Computer Vision and Pattern Recognition
Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D inconsistencies, e.g., objects and structures deforming under changes in camera pose, which can undermine user experience and simulation fidelity. Motivated by recent findings on representation alignment for diffusion models, we hypothesize that improving the multi-view consistency of video diffusion representations will yield more 3D-consistent video generation. Through detailed analysis on multiple recent camera-controlled video diffusion models we reveal strong correlations between 3D-consistent representations and videos. We also propose ViCoDR, a new approach for improving the 3D consistency of video models by learning multi-view consistent diffusion representations. We evaluate ViCoDR on camera controlled image-to-video, text-to-video, and multi-view generation models, demonstrating significant improvements in the 3D consistency of the generated videos. Project page: https://danier97.github.io/ViCoDR.
title View-Consistent Diffusion Representations for 3D-Consistent Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18991