Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Desai, Li, Jiahao, Tan, Hao, Sun, Xin, Shu, Zhixin, Zhou, Yi, Bi, Sai, Pirk, Sören, Kaufman, Arie E.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913308137750528
author Xie, Desai
Li, Jiahao
Tan, Hao
Sun, Xin
Shu, Zhixin
Zhou, Yi
Bi, Sai
Pirk, Sören
Kaufman, Arie E.
author_facet Xie, Desai
Li, Jiahao
Tan, Hao
Sun, Xin
Shu, Zhixin
Zhou, Yi
Bi, Sai
Pirk, Sören
Kaufman, Arie E.
contents Multi-view diffusion models, obtained by applying Supervised Finetuning (SFT) to text-to-image diffusion models, have driven recent breakthroughs in text-to-3D research. However, due to the limited size and quality of existing 3D datasets, they still suffer from multi-view inconsistencies and Neural Radiance Field (NeRF) reconstruction artifacts. We argue that multi-view diffusion models can benefit from further Reinforcement Learning Finetuning (RLFT), which allows models to learn from the data generated by themselves and improve beyond their dataset limitations during SFT. To this end, we introduce Carve3D, an improved RLFT algorithm coupled with a novel Multi-view Reconstruction Consistency (MRC) metric, to enhance the consistency of multi-view diffusion models. To measure the MRC metric on a set of multi-view images, we compare them with their corresponding NeRF renderings at the same camera viewpoints. The resulting model, which we denote as Carve3DM, demonstrates superior multi-view consistency and NeRF reconstruction quality than existing models. Our results suggest that pairing SFT with Carve3D's RLFT is essential for developing multi-view-consistent diffusion models, mirroring the standard Large Language Model (LLM) alignment pipeline. Our code, training and testing data, and video results are available at: https://desaixie.github.io/carve-3d.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13980
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning
Xie, Desai
Li, Jiahao
Tan, Hao
Sun, Xin
Shu, Zhixin
Zhou, Yi
Bi, Sai
Pirk, Sören
Kaufman, Arie E.
Computer Vision and Pattern Recognition
Machine Learning
Multi-view diffusion models, obtained by applying Supervised Finetuning (SFT) to text-to-image diffusion models, have driven recent breakthroughs in text-to-3D research. However, due to the limited size and quality of existing 3D datasets, they still suffer from multi-view inconsistencies and Neural Radiance Field (NeRF) reconstruction artifacts. We argue that multi-view diffusion models can benefit from further Reinforcement Learning Finetuning (RLFT), which allows models to learn from the data generated by themselves and improve beyond their dataset limitations during SFT. To this end, we introduce Carve3D, an improved RLFT algorithm coupled with a novel Multi-view Reconstruction Consistency (MRC) metric, to enhance the consistency of multi-view diffusion models. To measure the MRC metric on a set of multi-view images, we compare them with their corresponding NeRF renderings at the same camera viewpoints. The resulting model, which we denote as Carve3DM, demonstrates superior multi-view consistency and NeRF reconstruction quality than existing models. Our results suggest that pairing SFT with Carve3D's RLFT is essential for developing multi-view-consistent diffusion models, mirroring the standard Large Language Model (LLM) alignment pipeline. Our code, training and testing data, and video results are available at: https://desaixie.github.io/carve-3d.
title Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2312.13980