Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xinran, Diao, Muxi, Liu, Yuanzhi, Wang, Chunyu, Liang, Kongming, Ma, Zhanyu, Guo, Jun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918028225019904
author Wang, Xinran
Diao, Muxi
Liu, Yuanzhi
Wang, Chunyu
Liang, Kongming
Ma, Zhanyu
Guo, Jun
author_facet Wang, Xinran
Diao, Muxi
Liu, Yuanzhi
Wang, Chunyu
Liang, Kongming
Ma, Zhanyu
Guo, Jun
contents Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length to represent the detailness of the caption in the T2I training set. In this paper, we propose a new metric to estimate caption detailness based on two aspects: image coverage rate (ICR), which evaluates whether the caption covers all regions/objects in the image, and average object detailness (AOD), which quantifies the detailness of each object's description. Through experiments on the COCO dataset using ShareGPT4V captions, we demonstrate that T2I models trained on high-ICR and -AOD captions achieve superior performance on DPG and other benchmarks. Notably, our metric enables more effective data selection-training on only 20% of full data surpasses both full-dataset training and length-based selection method, improving alignment and reconstruction ability. These findings highlight the critical role of detail-aware metrics over length-based heuristics in caption selection for T2I tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
Wang, Xinran
Diao, Muxi
Liu, Yuanzhi
Wang, Chunyu
Liang, Kongming
Ma, Zhanyu
Guo, Jun
Computer Vision and Pattern Recognition
Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length to represent the detailness of the caption in the T2I training set. In this paper, we propose a new metric to estimate caption detailness based on two aspects: image coverage rate (ICR), which evaluates whether the caption covers all regions/objects in the image, and average object detailness (AOD), which quantifies the detailness of each object's description. Through experiments on the COCO dataset using ShareGPT4V captions, we demonstrate that T2I models trained on high-ICR and -AOD captions achieve superior performance on DPG and other benchmarks. Notably, our metric enables more effective data selection-training on only 20% of full data surpasses both full-dataset training and length-based selection method, improving alignment and reconstruction ability. These findings highlight the critical role of detail-aware metrics over length-based heuristics in caption selection for T2I tasks.
title Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.15172