Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jay Zhangjie, Zhang, Yuxuan, Turki, Haithem, Ren, Xuanchi, Gao, Jun, Shou, Mike Zheng, Fidler, Sanja, Gojcic, Zan, Ling, Huan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916641033420800
author Wu, Jay Zhangjie
Zhang, Yuxuan
Turki, Haithem
Ren, Xuanchi
Gao, Jun
Shou, Mike Zheng
Fidler, Sanja
Gojcic, Zan
Ling, Huan
author_facet Wu, Jay Zhangjie
Zhang, Yuxuan
Turki, Haithem
Ren, Xuanchi
Gao, Jun
Shou, Mike Zheng
Fidler, Sanja
Gojcic, Zan
Ling, Huan
contents Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a novel pipeline designed to enhance 3D reconstruction and novel-view synthesis through single-step diffusion models. At the core of our approach is Difix, a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by underconstrained regions of the 3D representation. Difix serves two critical roles in our pipeline. First, it is used during the reconstruction phase to clean up pseudo-training views that are rendered from the reconstruction and then distilled back into 3D. This greatly enhances underconstrained regions and improves the overall 3D representation quality. More importantly, Difix also acts as a neural enhancer during inference, effectively removing residual artifacts arising from imperfect 3D supervision and the limited capacity of current reconstruction models. Difix3D+ is a general solution, a single model compatible with both NeRF and 3DGS representations, and it achieves an average 2$\times$ improvement in FID score over baselines while maintaining 3D consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01774
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
Wu, Jay Zhangjie
Zhang, Yuxuan
Turki, Haithem
Ren, Xuanchi
Gao, Jun
Shou, Mike Zheng
Fidler, Sanja
Gojcic, Zan
Ling, Huan
Computer Vision and Pattern Recognition
Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a novel pipeline designed to enhance 3D reconstruction and novel-view synthesis through single-step diffusion models. At the core of our approach is Difix, a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by underconstrained regions of the 3D representation. Difix serves two critical roles in our pipeline. First, it is used during the reconstruction phase to clean up pseudo-training views that are rendered from the reconstruction and then distilled back into 3D. This greatly enhances underconstrained regions and improves the overall 3D representation quality. More importantly, Difix also acts as a neural enhancer during inference, effectively removing residual artifacts arising from imperfect 3D supervision and the limited capacity of current reconstruction models. Difix3D+ is a general solution, a single model compatible with both NeRF and 3DGS representations, and it achieves an average 2$\times$ improvement in FID score over baselines while maintaining 3D consistency.
title Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.01774