T2I-ConBench: Text-to-Image Benchmark for Continual Post-training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Zhehao, Liu, Yuhang, Lou, Yixin, He, Zhengbao, He, Mingzhen, Zhou, Wenxing, Li, Tao, Li, Kehan, Huang, Zeyi, Huang, Xiaolin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909620341047296
author Huang, Zhehao
Liu, Yuhang
Lou, Yixin
He, Zhengbao
He, Mingzhen
Zhou, Wenxing
Li, Tao
Li, Kehan
Huang, Zeyi
Huang, Xiaolin
author_facet Huang, Zhehao
Liu, Yuhang
Lou, Yixin
He, Zhengbao
He, Mingzhen
Zhou, Wenxing
Li, Tao
Li, Kehan
Huang, Zeyi
Huang, Xiaolin
contents Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of pretrained knowledge and undermines zero-shot compositionality. We observe that the absence of a standardized evaluation protocol hampers related research for continual post-training. To address this, we introduce T2I-ConBench, a unified benchmark for continual post-training of text-to-image models. T2I-ConBench focuses on two practical scenarios, item customization and domain enhancement, and analyzes four dimensions: (1) retention of generality, (2) target-task performance, (3) catastrophic forgetting, and (4) cross-task generalization. It combines automated metrics, human-preference modeling, and vision-language QA for comprehensive assessment. We benchmark ten representative methods across three realistic task sequences and find that no approach excels on all fronts. Even joint "oracle" training does not succeed for every task, and cross-task generalization remains unsolved. We release all datasets, code, and evaluation tools to accelerate research in continual post-training for text-to-image models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle T2I-ConBench: Text-to-Image Benchmark for Continual Post-training
Huang, Zhehao
Liu, Yuhang
Lou, Yixin
He, Zhengbao
He, Mingzhen
Zhou, Wenxing
Li, Tao
Li, Kehan
Huang, Zeyi
Huang, Xiaolin
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of pretrained knowledge and undermines zero-shot compositionality. We observe that the absence of a standardized evaluation protocol hampers related research for continual post-training. To address this, we introduce T2I-ConBench, a unified benchmark for continual post-training of text-to-image models. T2I-ConBench focuses on two practical scenarios, item customization and domain enhancement, and analyzes four dimensions: (1) retention of generality, (2) target-task performance, (3) catastrophic forgetting, and (4) cross-task generalization. It combines automated metrics, human-preference modeling, and vision-language QA for comprehensive assessment. We benchmark ten representative methods across three realistic task sequences and find that no approach excels on all fronts. Even joint "oracle" training does not succeed for every task, and cross-task generalization remains unsolved. We release all datasets, code, and evaluation tools to accelerate research in continual post-training for text-to-image models.
title T2I-ConBench: Text-to-Image Benchmark for Continual Post-training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.16875