Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Zeyi, Wu, Tong, Zhang, Pan, Zang, Yuhang, Dong, Xiaoyi, Xiong, Yuanjun, Lin, Dahua, Wang, Jiaqi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914963181797376
author Sun, Zeyi
Wu, Tong
Zhang, Pan
Zang, Yuhang
Dong, Xiaoyi
Xiong, Yuanjun
Lin, Dahua
Wang, Jiaqi
author_facet Sun, Zeyi
Wu, Tong
Zhang, Pan
Zang, Yuhang
Dong, Xiaoyi
Xiong, Yuanjun
Lin, Dahua
Wang, Jiaqi
contents Recent years have witnessed remarkable progress in multi-view diffusion models for 3D content creation. However, there remains a significant gap in image quality and prompt-following ability compared to 2D diffusion models. A critical bottleneck is the scarcity of high-quality 3D objects with detailed captions. To address this challenge, we propose Bootstrap3D, a novel framework that automatically generates an arbitrary quantity of multi-view images to assist in training multi-view diffusion models. Specifically, we introduce a data generation pipeline that employs (1) 2D and video diffusion models to generate multi-view images based on constructed text prompts, and (2) our fine-tuned 3D-aware MV-LLaVA for filtering high-quality data and rewriting inaccurate captions. Leveraging this pipeline, we have generated 1 million high-quality synthetic multi-view images with dense descriptive captions to address the shortage of high-quality 3D data. Furthermore, we present a Training Timestep Reschedule (TTR) strategy that leverages the denoising process to learn multi-view consistency while maintaining the original 2D diffusion prior. Extensive experiments demonstrate that Bootstrap3D can generate high-quality multi-view images with superior aesthetic quality, image-text alignment, and maintained view consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00093
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
Sun, Zeyi
Wu, Tong
Zhang, Pan
Zang, Yuhang
Dong, Xiaoyi
Xiong, Yuanjun
Lin, Dahua
Wang, Jiaqi
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Multimedia
Recent years have witnessed remarkable progress in multi-view diffusion models for 3D content creation. However, there remains a significant gap in image quality and prompt-following ability compared to 2D diffusion models. A critical bottleneck is the scarcity of high-quality 3D objects with detailed captions. To address this challenge, we propose Bootstrap3D, a novel framework that automatically generates an arbitrary quantity of multi-view images to assist in training multi-view diffusion models. Specifically, we introduce a data generation pipeline that employs (1) 2D and video diffusion models to generate multi-view images based on constructed text prompts, and (2) our fine-tuned 3D-aware MV-LLaVA for filtering high-quality data and rewriting inaccurate captions. Leveraging this pipeline, we have generated 1 million high-quality synthetic multi-view images with dense descriptive captions to address the shortage of high-quality 3D data. Furthermore, we present a Training Timestep Reschedule (TTR) strategy that leverages the denoising process to learn multi-view consistency while maintaining the original 2D diffusion prior. Extensive experiments demonstrate that Bootstrap3D can generate high-quality multi-view images with superior aesthetic quality, image-text alignment, and maintained view consistency.
title Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Multimedia
url https://arxiv.org/abs/2406.00093