MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Haofeng, Zhou, Yang, Wang, Ziheng, Xu, Zhengbo, Peng, Zhan, Ma, Jie, Liang, Jun, He, Shengfeng, Li, Jing
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910213615910912
author Liu, Haofeng
Zhou, Yang
Wang, Ziheng
Xu, Zhengbo
Peng, Zhan
Ma, Jie
Liang, Jun
He, Shengfeng
Li, Jing
author_facet Liu, Haofeng
Zhou, Yang
Wang, Ziheng
Xu, Zhengbo
Peng, Zhan
Ma, Jie
Liang, Jun
He, Shengfeng
Li, Jing
contents Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors offer visual fidelity but lack geometric correspondence. Existing methods either propagate geometric errors throughout generation or suffer from signal conflicts when fusing both statically. We introduce MoCam, which employs structured denoising dynamics to orchestrate a coordinated progression from geometry to appearance within the diffusion process. MoCam first leverages geometric priors in early stages to anchor coarse structures and tolerate their incompleteness, then switches to appearance priors in later stages to actively correct geometric errors and refine details. This design naturally unifies static and dynamic view synthesis by temporally decoupling geometric alignment and appearance refinement within the diffusion process. Experiments demonstrate that MoCam significantly outperforms prior methods, particularly when point clouds contain severe holes or distortions, achieving robust geometry-appearance disentanglement.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12119
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
Liu, Haofeng
Zhou, Yang
Wang, Ziheng
Xu, Zhengbo
Peng, Zhan
Ma, Jie
Liang, Jun
He, Shengfeng
Li, Jing
Computer Vision and Pattern Recognition
Graphics
Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors offer visual fidelity but lack geometric correspondence. Existing methods either propagate geometric errors throughout generation or suffer from signal conflicts when fusing both statically. We introduce MoCam, which employs structured denoising dynamics to orchestrate a coordinated progression from geometry to appearance within the diffusion process. MoCam first leverages geometric priors in early stages to anchor coarse structures and tolerate their incompleteness, then switches to appearance priors in later stages to actively correct geometric errors and refine details. This design naturally unifies static and dynamic view synthesis by temporally decoupling geometric alignment and appearance refinement within the diffusion process. Experiments demonstrate that MoCam significantly outperforms prior methods, particularly when point clouds contain severe holes or distortions, achieving robust geometry-appearance disentanglement.
title MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2605.12119