Pointmap-Conditioned Diffusion for Consistent Novel View Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thang-Anh-Quan, Piasco, Nathan, Roldão, Luis, Bennehar, Moussab, Tsishkou, Dzmitry, Caraffa, Laurent, Tarel, Jean-Philippe, Brémond, Roland
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912786993381376
author Nguyen, Thang-Anh-Quan
Piasco, Nathan
Roldão, Luis
Bennehar, Moussab
Tsishkou, Dzmitry
Caraffa, Laurent
Tarel, Jean-Philippe
Brémond, Roland
author_facet Nguyen, Thang-Anh-Quan
Piasco, Nathan
Roldão, Luis
Bennehar, Moussab
Tsishkou, Dzmitry
Caraffa, Laurent
Tarel, Jean-Philippe
Brémond, Roland
contents Synthesizing extrapolated views remains a difficult task, especially in urban driving scenes, where the only reliable sources of data are limited RGB captures and sparse LiDAR points. To address this problem, we present PointmapDiff, a framework for novel view synthesis that utilizes pre-trained 2D diffusion models. Our method leverages point maps (i.e., rasterized 3D scene coordinates) as a conditioning signal, capturing geometric and photometric priors from the reference images to guide the image generation process. With the proposed reference attention layers and ControlNet for point map features, PointmapDiff can generate accurate and consistent results across varying viewpoints while respecting geometric fidelity. Experiments on real-life driving data demonstrate that our method achieves high-quality generation with flexibility over point map conditioning signals (e.g., dense depth map or even sparse LiDAR points) and can be used to distill to 3D representations such as 3D Gaussian Splatting for improving view extrapolation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pointmap-Conditioned Diffusion for Consistent Novel View Synthesis
Nguyen, Thang-Anh-Quan
Piasco, Nathan
Roldão, Luis
Bennehar, Moussab
Tsishkou, Dzmitry
Caraffa, Laurent
Tarel, Jean-Philippe
Brémond, Roland
Computer Vision and Pattern Recognition
Synthesizing extrapolated views remains a difficult task, especially in urban driving scenes, where the only reliable sources of data are limited RGB captures and sparse LiDAR points. To address this problem, we present PointmapDiff, a framework for novel view synthesis that utilizes pre-trained 2D diffusion models. Our method leverages point maps (i.e., rasterized 3D scene coordinates) as a conditioning signal, capturing geometric and photometric priors from the reference images to guide the image generation process. With the proposed reference attention layers and ControlNet for point map features, PointmapDiff can generate accurate and consistent results across varying viewpoints while respecting geometric fidelity. Experiments on real-life driving data demonstrate that our method achieves high-quality generation with flexibility over point map conditioning signals (e.g., dense depth map or even sparse LiDAR points) and can be used to distill to 3D representations such as 3D Gaussian Splatting for improving view extrapolation.
title Pointmap-Conditioned Diffusion for Consistent Novel View Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.02913