MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sinha, Sankalp, Khan, Mohammad Sadil, Usama, Muhammad, Sam, Shino, Stricker, Didier, Ali, Sk Aziz, Afzal, Muhammad Zeshan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913758746509312
author Sinha, Sankalp
Khan, Mohammad Sadil
Usama, Muhammad
Sam, Shino
Stricker, Didier
Ali, Sk Aziz
Afzal, Muhammad Zeshan
author_facet Sinha, Sankalp
Khan, Mohammad Sadil
Usama, Muhammad
Sam, Shino
Stricker, Didier
Ali, Sk Aziz
Afzal, Muhammad Zeshan
contents Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 million 3D assets aggregated from seven major 3D datasets. Our contribution is a novel multi-stage annotation pipeline that integrates open-source pretrained multi-view VLMs and LLMs to automatically produce multi-level descriptions, ranging from detailed (150-200 words) to concise semantic tags (10-20 words). This structure supports both fine-grained 3D reconstruction and rapid prototyping. Furthermore, we incorporate human metadata from source datasets into our annotation pipeline to add domain-specific information in our annotation and reduce VLM hallucinations. Additionally, we develop MARVEL-FX3D, a two-stage text-to-3D pipeline. We fine-tune Stable Diffusion with our annotations and use a pretrained image-to-3D network to generate 3D textured meshes within 15s. Extensive evaluations show that MARVEL-40M+ significantly outperforms existing datasets in annotation quality and linguistic diversity, achieving win rates of 72.41% by GPT-4 and 73.40% by human evaluators. Project page is available at https://sankalpsinha-cmos.github.io/MARVEL/.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17945
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
Sinha, Sankalp
Khan, Mohammad Sadil
Usama, Muhammad
Sam, Shino
Stricker, Didier
Ali, Sk Aziz
Afzal, Muhammad Zeshan
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 million 3D assets aggregated from seven major 3D datasets. Our contribution is a novel multi-stage annotation pipeline that integrates open-source pretrained multi-view VLMs and LLMs to automatically produce multi-level descriptions, ranging from detailed (150-200 words) to concise semantic tags (10-20 words). This structure supports both fine-grained 3D reconstruction and rapid prototyping. Furthermore, we incorporate human metadata from source datasets into our annotation pipeline to add domain-specific information in our annotation and reduce VLM hallucinations. Additionally, we develop MARVEL-FX3D, a two-stage text-to-3D pipeline. We fine-tune Stable Diffusion with our annotations and use a pretrained image-to-3D network to generate 3D textured meshes within 15s. Extensive evaluations show that MARVEL-40M+ significantly outperforms existing datasets in annotation quality and linguistic diversity, achieving win rates of 72.41% by GPT-4 and 73.40% by human evaluators. Project page is available at https://sankalpsinha-cmos.github.io/MARVEL/.
title MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2411.17945