Any-to-3D Generation via Hybrid Diffusion Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Yijun, Ma, Yiwei, Ji, Jiayi, Sun, Xiaoshuai, Ji, Rongrong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929600425099264
author Fan, Yijun
Ma, Yiwei
Ji, Jiayi
Sun, Xiaoshuai
Ji, Rongrong
author_facet Fan, Yijun
Ma, Yiwei
Ji, Jiayi
Sun, Xiaoshuai
Ji, Rongrong
contents Recent progress in 3D object generation has been fueled by the strong priors offered by diffusion models. However, existing models are tailored to specific tasks, accommodating only one modality at a time and necessitating retraining to change modalities. Given an image-to-3D model and a text prompt, a naive approach is to convert text prompts to images and then use the image-to-3D model for generation. This approach is both time-consuming and labor-intensive, resulting in unavoidable information loss during modality conversion. To address this, we introduce XBind, a unified framework for any-to-3D generation using cross-modal pre-alignment techniques. XBind integrates an multimodal-aligned encoder with pre-trained diffusion models to generate 3D objects from any modalities, including text, images, and audio. We subsequently present a novel loss function, termed Modality Similarity (MS) Loss, which aligns the embeddings of the modality prompts and the rendered images, facilitating improved alignment of the 3D objects with multiple modalities. Additionally, Hybrid Diffusion Supervision combined with a Three-Phase Optimization process improves the quality of the generated 3D objects. Extensive experiments showcase XBind's broad generation capabilities in any-to-3D scenarios. To our knowledge, this is the first method to generate 3D objects from any modality prompts. Project page: https://zeroooooooow1440.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14715
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Any-to-3D Generation via Hybrid Diffusion Supervision
Fan, Yijun
Ma, Yiwei
Ji, Jiayi
Sun, Xiaoshuai
Ji, Rongrong
Computer Vision and Pattern Recognition
Recent progress in 3D object generation has been fueled by the strong priors offered by diffusion models. However, existing models are tailored to specific tasks, accommodating only one modality at a time and necessitating retraining to change modalities. Given an image-to-3D model and a text prompt, a naive approach is to convert text prompts to images and then use the image-to-3D model for generation. This approach is both time-consuming and labor-intensive, resulting in unavoidable information loss during modality conversion. To address this, we introduce XBind, a unified framework for any-to-3D generation using cross-modal pre-alignment techniques. XBind integrates an multimodal-aligned encoder with pre-trained diffusion models to generate 3D objects from any modalities, including text, images, and audio. We subsequently present a novel loss function, termed Modality Similarity (MS) Loss, which aligns the embeddings of the modality prompts and the rendered images, facilitating improved alignment of the 3D objects with multiple modalities. Additionally, Hybrid Diffusion Supervision combined with a Three-Phase Optimization process improves the quality of the generated 3D objects. Extensive experiments showcase XBind's broad generation capabilities in any-to-3D scenarios. To our knowledge, this is the first method to generate 3D objects from any modality prompts. Project page: https://zeroooooooow1440.github.io/.
title Any-to-3D Generation via Hybrid Diffusion Supervision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.14715