Saved in:
Bibliographic Details
Main Authors: Fu, Kang, Duan, Huiyu, Zhang, Zicheng, Liu, Xiaohong, Min, Xiongkuo, Wang, Jia, Zhai, Guangtao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.16915
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913704872771584
author Fu, Kang
Duan, Huiyu
Zhang, Zicheng
Liu, Xiaohong
Min, Xiongkuo
Wang, Jia
Zhai, Guangtao
author_facet Fu, Kang
Duan, Huiyu
Zhang, Zicheng
Liu, Xiaohong
Min, Xiongkuo
Wang, Jia
Zhai, Guangtao
contents Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and Model
Fu, Kang
Duan, Huiyu
Zhang, Zicheng
Liu, Xiaohong
Min, Xiongkuo
Wang, Jia
Zhai, Guangtao
Computer Vision and Pattern Recognition
Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively.
title Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.16915