Art3D: Training-Free 3D Generation from Flat-Colored Illustration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cong, Xiaoyan, Shen, Jiayi, Li, Zekun, Fu, Rao, Lu, Tao, Sridhar, Srinath
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915941024006144
author Cong, Xiaoyan
Shen, Jiayi
Li, Zekun
Fu, Rao
Lu, Tao
Sridhar, Srinath
author_facet Cong, Xiaoyan
Shen, Jiayi
Li, Zekun
Fu, Rao
Lu, Tao
Sridhar, Srinath
contents Large-scale pre-trained image-to-3D generative models have exhibited remarkable capabilities in diverse shape generations. However, most of them struggle to synthesize plausible 3D assets when the reference image is flat-colored like hand drawings due to the lack of 3D illusion, which are often the most user-friendly input modalities in art content creation. To this end, we propose Art3D, a training-free method that can lift flat-colored 2D designs into 3D. By leveraging structural and semantic features with pre-trained 2D image generation models and a VLM-based realism evaluation, Art3D successfully enhances the three-dimensional illusion in reference images, thus simplifying the process of generating 3D from 2D, and proves adaptable to a wide range of painting styles. To benchmark the generalization performance of existing image-to-3D models on flat-colored images without 3D feeling, we collect a new dataset, Flat-2D, with over 100 samples. Experimental results demonstrate the performance and robustness of Art3D, exhibiting superior generalizable capacity and promising practical applicability. Our source code and dataset will be publicly available on our project page: https://joy-jy11.github.io/ .
format Preprint
id arxiv_https___arxiv_org_abs_2504_10466
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Art3D: Training-Free 3D Generation from Flat-Colored Illustration
Cong, Xiaoyan
Shen, Jiayi
Li, Zekun
Fu, Rao
Lu, Tao
Sridhar, Srinath
Computer Vision and Pattern Recognition
Large-scale pre-trained image-to-3D generative models have exhibited remarkable capabilities in diverse shape generations. However, most of them struggle to synthesize plausible 3D assets when the reference image is flat-colored like hand drawings due to the lack of 3D illusion, which are often the most user-friendly input modalities in art content creation. To this end, we propose Art3D, a training-free method that can lift flat-colored 2D designs into 3D. By leveraging structural and semantic features with pre-trained 2D image generation models and a VLM-based realism evaluation, Art3D successfully enhances the three-dimensional illusion in reference images, thus simplifying the process of generating 3D from 2D, and proves adaptable to a wide range of painting styles. To benchmark the generalization performance of existing image-to-3D models on flat-colored images without 3D feeling, we collect a new dataset, Flat-2D, with over 100 samples. Experimental results demonstrate the performance and robustness of Art3D, exhibiting superior generalizable capacity and promising practical applicability. Our source code and dataset will be publicly available on our project page: https://joy-jy11.github.io/ .
title Art3D: Training-Free 3D Generation from Flat-Colored Illustration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.10466